Artificial Intelligence #ai#artificial intelligence
TxBench-PP Benchmark Reveals No AI Agent Reliably Matches Preclinical Pharmacology Decisions
Researchers introduced TxBench-PP, a verifiable benchmark for AI agents in small-molecule preclinical pharmacology. Across 16 model-harness configurations and 4,800 trajectories, no system reliably recovered pharmacology decisions; the top performer, Claude Opus 4.8 / Pi, passed 59.3% of endpoint attempts.
Jun 22, 2026 1 source