Static leaderboards that rank LLM agents by aggregate scores fail to predict real-world performance, according to a new paper on arXiv that proposes a shift to 'predictive validity'—measuring how well in-sample rankings correlate with out-of-distribution performance. The paper, authored by a large team including Dhaval C. Patel and Kaoutar El Maghraoui, consolidates the largest coordinated deep-dive of a single MCP-based industrial-agent benchmark to date, covering fourteen parallel implementation studies.
The Problem with Aggregate Scores
The paper argues that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings. The authors note that "recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability." Furthermore, "no single benchmark touches more than four or five of the dimensions that deployment exposes."
| Current Approach | Proposed Approach |
|---|---|
| Aggregate-score leaderboards | Predictive validity ranking |
| In-sample mean | Correlation between in-sample and out-of-sample rank |
| Rank instability documented | Falsifiable out-of-distribution criteria |
The authors consolidate those fourteen studies with seven prior agent benchmarks to support their critique.
Introducing Predictive Validity
The authors propose ranking configurations by predictive validity, defined as "the correlation between in-sample and out-of-sample rank, rather than in-sample mean." They report a twelve-tier measurement apparatus that "exposes the deployment-relevant dimensions HELM and its agent-era successors collapse." HELM (Holistic Evaluation of Language Models) is a well-known framework; the paper suggests it and similar tools miss key dimensions relevant to deployed agents.
A Twelve-Tier Measurement Apparatus
The twelve-tier apparatus operationalizes the position through three falsifiable out-of-distribution criteria with explicit thresholds. The paper states that "existing evidence partly supports it but is too thin to confirm." The authors close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report. The studies included new asset classes, multi-modal visual extensions, alternative orchestrations, retrieval strategies, reasoning modes, and infrastructure optimizations.
Implications for Enterprise AI Procurement
For CTOs and technology procurement leaders evaluating LLM agents, this research underscores the risk of relying on static leaderboard scores. The paper suggests that enterprise buyers should demand out-of-distribution validation and metrics that reflect deployment diversity. The twelve-tier apparatus could serve as a template for more robust evaluation, though the authors acknowledge the evidence is not yet conclusive. As the field advances, predictive validity may become a standard requirement for AI vendor assessments.