iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› Llms ›› New MBABench Evaluates LLM Agents on End-to-End Finance Spreadsheet Tasks

New MBABench Evaluates LLM Agents on End-to-End Finance Spreadsheet Tasks

MBABench, a new benchmark from researchers, evaluates LLM agents on end-to-end spreadsheet tasks in finance, focusing on modeling and scenario analysis. The benchmark assesses accuracy, formula use, and formatting. Claude family models lead but still fall short of professional standards.

iG
iGEN Editorial
June 16, 2026
New MBABench Evaluates LLM Agents on End-to-End Finance Spreadsheet Tasks

Enterprise finance teams routinely build spreadsheets for modeling, forecasting, and scenario analysis. Yet, as a new benchmark reveals, current LLM agents are not yet capable of reliably producing professional-quality spreadsheets from scratch.

According to a paper published on arXiv titled "MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance," researchers including Yen, Thomson, Poeltl, and others developed one of the first evaluations of agents on complete spreadsheet workflows. The benchmark addresses a gap where existing spreadsheet benchmarks focus only on question-answering or single-formula edits, not on end-to-end artifact creation.

The Benchmark Design

MBABench targets economically critical financial workflows such as financial modeling, forecasting, and scenario analysis. Recognising that spreadsheet deliverables are routinely reviewed by multiple stakeholders, the researchers designed an evaluation taxonomy with three dimensions: Accuracy, Formula, and Format. Each dimension comprises fine-grained criteria reflecting professional standards.

The tasks require agents to produce entire spreadsheets from high-level user instructions, mimicking real-world demands where a deliverable must be readable, accurate, and easy to modify.

Key Findings: Claude Leads, But All Fall Short

In the evaluation, the Claude family of models led the benchmark and produced the most professional-looking outputs in a qualitative review. However, even the strongest agents frequently fell short of professional finance standards. Performance degraded sharply as task difficulty increased beyond a few chained calculations.

The paper states that "current agents are not yet able to reliably produce professional-quality spreadsheets at the level of complexity real-world workflows demand." This suggests that while progress has been made, enterprise-grade automation of spreadsheet creation remains out of reach.

Implications for Enterprise Technology Leaders

For CTOs and digital transformation leaders in finance and adjacent sectors like supply chain, the findings highlight a maturity gap. While LLM agents can handle simple formula tasks, they struggle with the multidimensional requirements of professional financial modeling. The benchmark's emphasis on readability and ease of modification underscores that enterprise users expect not just correct outputs, but outputs that can be reviewed and iterated upon by teams.

As frontier AI labs continue to develop agents for end-to-end workflows, MBABench provides a structured way to measure progress. The paper is available on arXiv under a Creative Commons license.

In summary, the Claude family leads but no current agent meets professional standards, especially as complexity increases. This benchmark sets a new bar for evaluating whether AI can truly replace or augment human spreadsheet work in finance.


Sources:

Keep Reading

Recommended Stories

Is AI facing a big financial reckoning? Chip stocks tumble as investor euphoria fades Technology

Is AI facing a big financial reckoning? Chip stocks tumble as investor euphoria fades

Sharp falls in chip maker stocks have triggered investor concerns that the AI euphoria is fading. Korean chip makers SK Hynix and Samsung dropped 46% and 35% in a month, while US firms Micron and Intel fell 28% and 35%. China's reported chip breakthrough and worries over AI monetization add pressure. Despite the sell-off, Eileen Burbidge calls it a healthy correction after massive gains.

July 29, 2026
AI Is Coming for Accounts Receivable’s Busywork, Not Its Jobs, Says FreightTech CEO Technology

AI Is Coming for Accounts Receivable’s Busywork, Not Its Jobs, Says FreightTech CEO

Stuut CEO Tarek Alaruri discusses how AI automates accounts receivable in freight brokerage, cutting overdue invoices by 40% by unifying previously siloed processes across the order-to-cash lifecycle.

July 22, 2026
Benchmarking Agentic Review Systems: AI Peer Review Achieves 83% Pairwise Accuracy but Falls Short on Error Detection Technology

Benchmarking Agentic Review Systems: AI Peer Review Achieves 83% Pairwise Accuracy but Falls Short on Error Detection

A study by Nguyen et al. benchmarks two open-source and one proprietary AI review system on peer review tasks. The best configuration (OpenAIReview + GPT-5.5) achieves 83.0% pairwise accuracy in tracking paper quality but only 71.6% recall in detecting injected errors. User feedback shows a positive-to-negative vote ratio of 1.44:1, with common complaints about false positives. The research highlights both the potential and limitations of current AI agents in evaluation tasks.

July 8, 2026
TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement Technology

TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement

A new benchmark called TERMS-Bench goes beyond deal rate to diagnose why LLM negotiation agents fail, evaluating 13 frontier models on surplus extraction, cue use, belief calibration, and compliance. For enterprise procurement and trade, this offers actionable insights into AI agent weaknesses.

June 17, 2026