Topic
methodology
Bayesian Inference and Decision Audits Reveal Unreliability in Frontier AI Evaluation Archives
A new arXiv paper by Yanan Long applies Bayesian inference and decision audits to public archives of frontier AI evaluations, revealing that terminal leaderboard interpretations can be misleading due to selective time series, reporting rules, and missingness. The study examines archives including LiveBench, Open LLM Leaderboard v2, LMArena, GAIA, and tau-bench, and finds that a candidate selection-aware frontier model fails synthetic recovery and uncertainty calibration. The proposed archive-and-adjudication protocol reconstructs histories and falsifies unsupported claims.
MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance
Researchers propose MA-SBI, a misspecification-aware simulation-based inference framework that leverages unstructured side-channel information—such as regime labels or policy bulletins—to correct posterior estimates without requiring ground-truth parameter pairs. The method matches oracle performance on hide-the-calibration benchmarks and improves log-likelihood on real COVID epidemiological data.
India Does Not Use Methodology Changes to Inflate Growth Numbers, CEA Nageswaran Defends GDP Data
Chief Economic Adviser V Anantha Nageswaran has defended India's GDP data, stating that the country does not use methodology or base-year changes to artificially inflate economic growth figures. In an ANI interview, he said GDP is an estimate, and India actually lowered its GDP figure after a recent rebasing exercise, contrary to many other nations. He also noted that international institutions like the IMF have only questioned methodology, not reliability, and that criticism often stems from unmet expectations.