iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Taiwan Charts Offshore Wind Growth to 18 GW by 2039 in New Energy Roadmap Low Ending Stocks Will Likely Force India to Stop Sugar Export, Ethanol Diversion Early Next Season India Has Ingredients to Become Alternative Protein Manufacturing Hub, Says GFI India MD Sneha Singh UP and NS CEOs Say Latest Rail Merger Filing Additions Further Enhance Competition Greek Owner JME Navigation Adds Another Ultramax to New Dayang Orderbook, Securing 2030 Delivery Slot C.H. Robinson Faces $604 Million Verdict: Vicarious Liability and Negligent Hiring Reshape Broker Risk After Deadly Crash ArcBest's Q2 Results Show Operational Recovery in LTL and Brokerage Segments South Africa's Prince Edward Graving Dock Sees Major Upgrade with Troy Docking, Boosting Ship Repair Options India Salary Hikes Projected at 8.6%-10.2%; EV, Fintech, Healthcare Lead Pay Gains Carrier diversification unravels the last-mile delivery duopoly as shippers seek alternatives Taiwan Charts Offshore Wind Growth to 18 GW by 2039 in New Energy Roadmap Low Ending Stocks Will Likely Force India to Stop Sugar Export, Ethanol Diversion Early Next Season India Has Ingredients to Become Alternative Protein Manufacturing Hub, Says GFI India MD Sneha Singh UP and NS CEOs Say Latest Rail Merger Filing Additions Further Enhance Competition Greek Owner JME Navigation Adds Another Ultramax to New Dayang Orderbook, Securing 2030 Delivery Slot C.H. Robinson Faces $604 Million Verdict: Vicarious Liability and Negligent Hiring Reshape Broker Risk After Deadly Crash ArcBest's Q2 Results Show Operational Recovery in LTL and Brokerage Segments South Africa's Prince Edward Graving Dock Sees Major Upgrade with Troy Docking, Boosting Ship Repair Options India Salary Hikes Projected at 8.6%-10.2%; EV, Fintech, Healthcare Lead Pay Gains Carrier diversification unravels the last-mile delivery duopoly as shippers seek alternatives
Home ›› Technology ›› Ai ›› CODA-BENCH: New Benchmark Reveals Code Agents Struggle with Data-Intensive Tasks

CODA-BENCH: New Benchmark Reveals Code Agents Struggle with Data-Intensive Tasks

A new benchmark called CODA-BENCH evaluates code agents on data-intensive tasks using a Kaggle-based sandbox. It comprises 1,009 tasks across 31 communities, each with an average of 980 files. Even top-performing agents achieve only a 61.1% success rate, highlighting a significant gap in integrating data discovery with code execution.

iG
iGEN Editorial
June 16, 2026
CODA-BENCH: New Benchmark Reveals Code Agents Struggle with Data-Intensive Tasks

Enterprise AI agents are increasingly expected to handle complex, data-heavy workflows, but current benchmarks often test code and data capabilities in isolation. A new preprint from researchers including Zhang, Yuxin, Fan, Ju, Meihao, Shaolei, Du, and Xiaoyong introduces CODA-BENCH, the first benchmark designed to jointly evaluate code intelligence and data intelligence in a realistic, data-intensive environment. The findings, posted on arXiv, reveal that even the most advanced code agents struggle to integrate data discovery with code generation, achieving only a 61.1% success rate.

Benchmark Design and Scope

CODA-BENCH is built around a Linux sandbox that mirrors the Kaggle ecosystem, containing hundreds of datasets. Agents must actively explore complex file hierarchies to identify relevant resources and then generate code for data-driven analytical tasks. The benchmark comprises 1,009 tasks spanning 31 communities, with each task environment containing an average of 980 files—simulating the scale and noise of real-world data environments.

The researchers explicitly designed CODA-BENCH to bridge the gap between existing benchmarks, which typically evaluate code-centric or data-centric abilities separately, and real development scenarios where both are required simultaneously.

Findings: A Significant Capability Gap

Evaluations of several advanced code agents on CODA-BENCH showed that even top-performing systems struggle effectively to combine data discovery with code execution. The highest success rate recorded was 61.1%, indicating a substantial shortfall in current agentic capabilities for data-intensive tasks. The results suggest that agents tend to perform well on either code generation or data retrieval, but fail to coordinate the two seamlessly.

Benchmark Metric Value
Total tasks 1,009
Communities represented 31
Average files per task 980
Top agent success rate 61.1%

The paper notes that these results "highlight a substantial gap in current agentic capabilities for data-intensive tasks and point to promising directions for future research."

Why This Matters for Enterprise AI

For enterprise technology decision-makers, CODA-BENCH serves as a realistic stress test for AI agents that might be deployed to handle data-heavy operations such as supply chain analytics, trade documentation processing, or financial data reconciliation. The benchmark's reliance on large file systems and heterogeneous data sources mirrors the complexity of corporate databases and data lakes. The 61.1% success ceiling suggests that current agent technology is not yet reliable enough for unsupervised deployment in data-intensive enterprise workflows without human oversight. Vendors developing code-generation tools and AI agents should prioritize integrating robust data-discovery capabilities alongside coding skills.

The researchers have made the benchmark available through arXiv, and the underlying code and data are expected to be released to the community, enabling further research and improvement of agent architectures. CODA-BENCH provides a clear metric for progress: until agents consistently score above 90% in such realistic settings, enterprises should approach autonomous data-task automation with caution.


Sources:

Keep Reading

Recommended Stories

ScholarQuest Benchmark Reveals Gaps in Agentic Academic Paper Search for Enterprise AI Technology

ScholarQuest Benchmark Reveals Gaps in Agentic Academic Paper Search for Enterprise AI

A new benchmark called ScholarQuest evaluates LLM-based agents for academic paper search. Built from over 1,000 computer science topics and four research intents, it provides scalable answer construction and a shared retrieval backend. Results show agentic methods beat single-shot retrieval but the top agent only achieves 0.314 Recall@100, indicating significant room for improvement in agentic search.

July 8, 2026
Controlled Benchmark Finds No Quantum Advantage in Brain MRI Data Augmentation Technology

Controlled Benchmark Finds No Quantum Advantage in Brain MRI Data Augmentation

A controlled benchmark study by Haider and Figini shows that quantum-latent GAN augmentation does not improve brain MRI classification over real-data-only training or classical GANs. The quantum and classical generators were statistically indistinguishable across all data fractions from 5% to 100%.

June 21, 2026
New PhysAssistBench Tests Medical LLMs on Interactive Doctor-Patient-EHR Coordination Technology

New PhysAssistBench Tests Medical LLMs on Interactive Doctor-Patient-EHR Coordination

Researchers introduce PhysAssistBench, a benchmark for evaluating medical LLMs on interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, it uses a scalable pipeline to create agentic patients. Experiments show leading LLMs remain unreliable, highlighting the need for coordination across knowledge, communication, and systems.

June 21, 2026
CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation Technology

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

CADBench is a unified benchmark for multimodal CAD program generation, containing 18,000 evaluation samples across six benchmark families, five input modalities, and six metrics. The benchmark evaluates eleven AI systems, generating over 1.4 million CAD programs, and reveals key failure modes in current approaches.

June 21, 2026