Artificial Intelligence #llm#a/b testing
LLM-Based A/B Testing Needs Calibration: New Statistical Framework Reveals 39% Accuracy Gap
A new paper from researchers at arXiv develops a statistical framework for using large language models (LLMs) as surrogates for human participants in A/B tests. The framework adapts surrogate endpoint theory, showing that raw LLM predictions recover only 39% of the human treatment effect, but calibration can close the gap. The study cautions that LLM-based A/B testing yields correct results only by assumption, whereas human testing is correct by design.
Jun 20, 2026 1 source