Artificial Intelligence #staminabench#coding agents
New StaminaBench Benchmark Reveals Coding Agents Fail After 5-6 Turns
Researchers introduce StaminaBench, a benchmark that measures how many consecutive interaction turns coding agents can handle. Testing six harnesses and seven open-source LLMs over 100-turn scenarios, they found all models fail within 5-6 turns. Providing test feedback improved passed turn count by up to 12x, highlighting the importance of iterative testing.
Jun 22, 2026 2 sources