iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› Llms ›› CoffeeBench: New Benchmark Evaluates LLM Agents in Multi-Agent Economic Simulations

CoffeeBench: New Benchmark Evaluates LLM Agents in Multi-Agent Economic Simulations

Researchers introduce CoffeeBench, a benchmark for evaluating LLM agents in a long-horizon multi-agent economy. The 90-day simulation features farmers, roasters, and retailers, with models controlling one roaster. All models outperformed a passive baseline, but Claude Haiku 4.5 showed an idle-drift failure mode.

iG
iGEN Editorial
June 16, 2026
CoffeeBench: New Benchmark Evaluates LLM Agents in Multi-Agent Economic Simulations

Evaluating large language model (LLM) agents in economic systems presents unique challenges not addressed by existing benchmarks. Most current evaluations test a single agent interacting with a passive environment, but real-world economies are multi-agent, requiring autonomous agents to communicate, negotiate, and transact over extended periods. To fill this gap, researchers from a team including Sugiura, Issa, Hattori, Daichi, Araragi, Kazuo, and others introduced CoffeeBench, a benchmark designed to assess LLM agents in a long-horizon multi-agent economy composed of heterogeneous firms, according to the paper published on arXiv.

How CoffeeBench Works

CoffeeBench simulates a coffee supply chain over 90 days. The simulation includes two farmers, two roasters, and two retailers, each operating autonomously. The objective for each firm is to maximize cumulative net income through communication and transactions while managing cash, inventory, and pricing. The evaluated LLM model controls one coffee roaster, while the remaining five firms are controlled by fixed reference agents. This setup tests an agent's ability to engage in sustained economic interaction, including negotiation and strategic planning.

Key Findings from the Evaluation

The researchers tested several recent open-weight and proprietary LLMs. According to the paper, all models outperformed a passive baseline that takes no actions, with most achieving positive net income. Analysis of agent behavior revealed substantial differences in long-horizon economic interaction. Higher-performing models communicated more actively with other firms. In contrast, Claude Haiku 4.5 exhibited an "idle-drift failure mode," repeatedly choosing inaction despite producing coherent assessments and plans. This finding highlights a critical gap between an agent's reasoning capabilities and its ability to execute economically productive actions.

Implications for Enterprise AI

While CoffeeBench is a research benchmark, its methodology has direct relevance for enterprise technology leaders evaluating AI for supply chain and logistics automation. The need for agents that can autonomously handle procurement, pricing, and inventory management over long horizons is growing. The benchmark provides a structured way to compare models on these capabilities, revealing that communication frequency and consistent execution are as important as raw reasoning power. The researchers have released the code and agent trajectories to support future research, enabling organizations to test their own models against the CoffeeBench environment.


Sources:

Keep Reading

Recommended Stories

LedgerAgent: A New Method for Policy-Adherent Tool-Calling AI Agents in Customer Service Technology

LedgerAgent: A New Method for Policy-Adherent Tool-Calling AI Agents in Customer Service

Researchers introduce LedgerAgent, an inference-time method that maintains observed task states in a separate ledger and checks policy constraints before tool calls, improving pass^k metrics across four customer-service domains. The approach addresses common failure modes where agents use stale or incorrect information or violate domain policies.

June 20, 2026
Hidden Anchors Reveal Why Multi-Agent LLM Deliberation Escapes Groupthink Technology

Hidden Anchors Reveal Why Multi-Agent LLM Deliberation Escapes Groupthink

A new paper from arXiv models multi-agent LLM deliberation as a closed-loop dynamical system where each agent has a hidden internal belief, or anchor, that continually pulls its opinion. The model explains how agents' confidence can climb past where any agent started, escaping the convex hull of initial beliefs. Tests across three open-weight model families show the anchor's influence is a spectrum.

June 20, 2026
AdaSTORM Breakthrough Scales LLM Reasoning to Thousand-Node Dynamic Graphs, Paves Way for Supply Chain AI Technology

AdaSTORM Breakthrough Scales LLM Reasoning to Thousand-Node Dynamic Graphs, Paves Way for Supply Chain AI

AdaSTORM, a new multi-agent AI framework, scales large language model reasoning to dynamic graphs of up to thousand nodes with over 90% accuracy. The approach uses adaptive partitioning and collaborative reasoning to overcome limitations of current LLMs, which can only handle tens of nodes. This breakthrough could enable AI-driven analysis of complex, evolving networks such as supply chains.

June 16, 2026
The Chatbot That Foretold Why People Share Secrets With ChatGPT Technology

The Chatbot That Foretold Why People Share Secrets With ChatGPT

A new book, 'Inventing ELIZA', recovers the source code of the 1960s chatbot from MIT Archives. The 'ELIZA effect' shows how people attribute empathy to computers, with profound implications for modern AI trust and enterprise deployment.

July 14, 2026