iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity US stock market flat as S&P 500 and Dow barely move, Nasdaq slides over 1% on chip rout Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity US stock market flat as S&P 500 and Dow barely move, Nasdaq slides over 1% on chip rout
Home ›› Technology ›› Ai ›› Llms ›› TxBench-PP Benchmark Reveals No AI Agent Reliably Matches Preclinical Pharmacology Decisions

TxBench-PP Benchmark Reveals No AI Agent Reliably Matches Preclinical Pharmacology Decisions

Researchers introduced TxBench-PP, a verifiable benchmark for AI agents in small-molecule preclinical pharmacology. Across 16 model-harness configurations and 4,800 trajectories, no system reliably recovered pharmacology decisions; the top performer, Claude Opus 4.8 / Pi, passed 59.3% of endpoint attempts.

iG
iGEN Editorial
June 22, 2026
TxBench-PP Benchmark Reveals No AI Agent Reliably Matches Preclinical Pharmacology Decisions

Artificial intelligence agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions. A new benchmark, TxBench-PP (TherapeuticsBench Preclinical Pharmacology), aims to provide that evaluation by testing whether agents can recover accurate conclusions from real-world assay data rather than memorized facts from literature.

Benchmark Design and Scope

According to the paper by Le, Hannah, Ramasamy, Ramesh, Urrutia, Alex, Yazdani, Mahsa, Proctor, Tim, Workman, and Kenny, posted on arXiv, TxBench-PP is the first focused slice of a broader TherapeuticsBench effort that will span drug-discovery stages and therapeutic modalities. The benchmark contains 100 evaluations indexed by program stage, assay type, and task structure. It covers five categories of reasoning:

  • Mechanism-of-action (MoA) and pharmacodynamic (PD) reasoning
  • Compound-target engagement
  • Causal target validation
  • Developability and safety
  • Translational efficacy

Agents receive realistic workflow snapshots, inspect files in a coding environment, and return structured answers that are graded deterministically. This design ensures that success depends on analytical capability rather than memorization.

Performance Results Across Models

The researchers tested 16 model-harness configurations, comprising 11 models and 4,800 trajectories. No system reliably recovered preclinical pharmacology decisions. The strongest configuration, Claude Opus 4.8 / Pi, passed 59.3% of endpoint attempts (178/300; 95% CI, 51.1-67.6). The second-best was GPT-5.5 / Pi at 55.3% (166/300; 47.0-63.6).

Model Configuration Pass Rate Endpoint Attempts 95% Confidence Interval
Claude Opus 4.8 / Pi 59.3% 178/300 51.1-67.6
GPT-5.5 / Pi 55.3% 166/300 47.0-63.6

All other configurations performed lower, though specific scores were not detailed in the paper. The results indicate that even the most advanced models fall short of reliable performance in this domain.

Implications for AI in Drug Discovery

The gap between current AI capabilities and the demands of preclinical pharmacology is significant. The benchmark's structure — requiring agents to work with realistic assay data and make reasoning chains — exposes weaknesses in models that excel at language tasks but struggle with data-intensive analytical workflows. For enterprise technology decision-makers, TxBench-PP demonstrates that deploying AI agents in high-stakes scientific environments requires rigorous, domain-specific validation. The benchmark provides a template for evaluating AI systems before trusting them with real drug development decisions. While the current models show promise, none meet the reliability threshold needed for unsupervised use in preclinical pharmacology. Future work on the broader TherapeuticsBench will likely continue to push models toward better performance.


Sources:

Keep Reading

Recommended Stories

Bombay High Court to Hear Gadkari's Suit Against Meta, X, Google Over Deepfakes on August 5 Technology

Bombay High Court to Hear Gadkari's Suit Against Meta, X, Google Over Deepfakes on August 5

The Bombay High Court has scheduled August 5 for hearing Union Minister Nitin Gadkari's civil suit against Meta, X Corp, Google LLC and unknown persons over defamatory deepfakes and AI-generated posts. Gadkari alleges the fake content falsely portrays him as personally responsible for the ethanol-blending programme and claims financial benefit to him and his family, causing irreparable harm to his reputation and personality rights.

July 28, 2026
Can the New York Times Save Journalism From Our AI Overlords? Technology

Can the New York Times Save Journalism From Our AI Overlords?

New York Times publisher A.G. Sulzberger discusses the $20 million copyright lawsuit against OpenAI and Microsoft, the Trump administration's attacks on press freedom, and the existential challenges facing journalism as AI reshapes information consumption.

July 28, 2026
Beijing Accuses US AI Firms of Using Chinese Models for Training Technology

Beijing Accuses US AI Firms of Using Chinese Models for Training

The Chinese commerce ministry accused US artificial intelligence firms of using Chinese models to train their own AI systems through a process called distillation. This comes after US Treasury Secretary Scott Bessent threatened sanctions against China over alleged technology theft. China defended distillation as a widely used industry practice and vowed to take all necessary measures to safeguard its interests.

July 28, 2026
Inside Trump's AI Brain Trust: The Officials Shaping US Tech Policy Technology

Inside Trump's AI Brain Trust: The Officials Shaping US Tech Policy

The Trump administration relies on a small, scattered group of officials to decide AI policy and export controls. Key figures include Commerce Secretary Howard Lutnick, acting CAISI chief Arvind Raman, National Cyber Director Sean Cairncross, and former AI czar David Sacks, each with differing views on regulating Chinese AI.

July 27, 2026