Artificial Intelligence #ai#artificial intelligence
Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation
A new paper shows that pass@k, the standard metric for estimating math-reasoning difficulty, has a blind spot: 10.3–22.9% of examples deemed impossible by sampling are actually solvable via activation grafting. The finding challenges current practices in RL training, data curation, and verifier design.
Jul 8, 2026 1 source