Artificial Intelligence #code agents#benchmark
CODA-BENCH: New Benchmark Reveals Code Agents Struggle with Data-Intensive Tasks
A new benchmark called CODA-BENCH evaluates code agents on data-intensive tasks using a Kaggle-based sandbox. It comprises 1,009 tasks across 31 communities, each with an average of 980 files. Even top-performing agents achieve only a 61.1% success rate, highlighting a significant gap in integrating data discovery with code execution.
Jun 16, 2026 1 source