iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Taiwan Charts Offshore Wind Growth to 18 GW by 2039 in New Energy Roadmap Low Ending Stocks Will Likely Force India to Stop Sugar Export, Ethanol Diversion Early Next Season India Has Ingredients to Become Alternative Protein Manufacturing Hub, Says GFI India MD Sneha Singh UP and NS CEOs Say Latest Rail Merger Filing Additions Further Enhance Competition Greek Owner JME Navigation Adds Another Ultramax to New Dayang Orderbook, Securing 2030 Delivery Slot C.H. Robinson Faces $604 Million Verdict: Vicarious Liability and Negligent Hiring Reshape Broker Risk After Deadly Crash ArcBest's Q2 Results Show Operational Recovery in LTL and Brokerage Segments South Africa's Prince Edward Graving Dock Sees Major Upgrade with Troy Docking, Boosting Ship Repair Options India Salary Hikes Projected at 8.6%-10.2%; EV, Fintech, Healthcare Lead Pay Gains Carrier diversification unravels the last-mile delivery duopoly as shippers seek alternatives Taiwan Charts Offshore Wind Growth to 18 GW by 2039 in New Energy Roadmap Low Ending Stocks Will Likely Force India to Stop Sugar Export, Ethanol Diversion Early Next Season India Has Ingredients to Become Alternative Protein Manufacturing Hub, Says GFI India MD Sneha Singh UP and NS CEOs Say Latest Rail Merger Filing Additions Further Enhance Competition Greek Owner JME Navigation Adds Another Ultramax to New Dayang Orderbook, Securing 2030 Delivery Slot C.H. Robinson Faces $604 Million Verdict: Vicarious Liability and Negligent Hiring Reshape Broker Risk After Deadly Crash ArcBest's Q2 Results Show Operational Recovery in LTL and Brokerage Segments South Africa's Prince Edward Graving Dock Sees Major Upgrade with Troy Docking, Boosting Ship Repair Options India Salary Hikes Projected at 8.6%-10.2%; EV, Fintech, Healthcare Lead Pay Gains Carrier diversification unravels the last-mile delivery duopoly as shippers seek alternatives
Home ›› Technology ›› Ai ›› Llms ›› IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

IHBench, a new benchmark from researchers including Salimi et al., evaluates how voice agents recover after interruptions in structured enterprise workflows. The benchmark tests 27 audio-language models from OpenAI, Google, and the open-weight community, finding that closed-weight models are consistently more robust, degrading 3.3x more slowly in long conversations.

iG
iGEN Editorial
June 22, 2026
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

Voice agents deployed in structured workflows—such as customer service, healthcare scheduling, and account management—must handle frequent user interruptions while maintaining progress through multi-step procedures. According to a new paper by Salimi, Ahmad, Ma, Wentao, Tang, Yuzhi, Shen, Dongming, Li, Mu, and Smola, Alex, titled "IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows" (arXiv 2606.19595), existing benchmarks for speech-capable models focus only on the timing of interruptions: barge-in detection, endpointing, and turn-taking dynamics. They leave unmeasured what happens after the interruption: does the agent resume the workflow at the correct step? Does it address the user's interjection? Does it avoid re-delivering content the user already heard?

IHBench: A Structured Evaluation Framework

The researchers introduce IHBench (Interruption Handling Benchmark), which evaluates post-interruption recovery in voice agents executing state-machine-driven workflows across 10 enterprise domains. The benchmark injects six interruption types at controlled points mid-utterance, with per-interruption evaluation rubrics generated alongside the data. Each interruption is scored on two axes: task fulfillment and recovery quality.

Key Findings: Closed-Weight Models Lead

The study evaluated 27 audio-language model configurations from OpenAI, Google, and the open-weight community. The results show that models vary widely, and recovery quality depends strongly on the interruption type. Across experiments, closed-weight models are consistently more robust to interruptions than open-weight ones. Specifically:

  • Closed-weight models win far more often on task fulfillment.
  • They degrade roughly 3.3x more slowly as conversations grow longer.
  • They show no audio-versus-text modality gap, whereas open-weight models lose ground on all three metrics.

A table summarising the key comparisons:

Metric Closed-Weight Models Open-Weight Models
Task fulfillment wins Far more often Fewer wins
Degradation rate (long conversations) 3.3x slower degradation Faster degradation
Audio vs. text modality gap No gap Gap present

Validation and Distinct Capability

A human study validated the LLM judge against human annotators, confirming the benchmark's reliability. A cross-benchmark analysis against AudioMultiChallenge indicates that recovery quality is a largely distinct capability axis, not fully captured by existing benchmarks.

Implications for Enterprise Adoption

For CTOs and digital transformation leaders evaluating voice agents for structured workflow domains—such as logistics customer service, healthcare scheduling, or account management—these findings highlight the importance of testing post-interruption recovery. The benchmark provides a standardised way to compare models from major providers like OpenAI and Google against open-weight alternatives. Closed-weight models currently offer superior robustness, especially in long interactions, which is critical for enterprise deployments where interruptions are common and continuity matters. The research also underscores that recovery quality is a separate axis from basic interruption detection, warranting dedicated evaluation in procurement decisions.


Sources:

Keep Reading

Recommended Stories

RTSGameBench Benchmark Tests Strategic Reasoning in Vision-Language Models Technology

RTSGameBench Benchmark Tests Strategic Reasoning in Vision-Language Models

A new benchmark called RTSGameBench evaluates strategic reasoning in vision-language models (VLMs) using the real-time strategy game Beyond All Reason. The benchmark includes diagnostic mini-games, diverse matchup structures, and a self-evolving generation framework. Initial tests show state-of-the-art VLMs struggle with tighter coordination, multiagent tasks, and increased scale.

June 21, 2026
Multi-LCB: New Benchmark Evaluates LLMs Across 12 Programming Languages Technology

Multi-LCB: New Benchmark Evaluates LLMs Across 12 Programming Languages

A new benchmark, Multi-LCB, extends the popular LiveCodeBench to 12 programming languages, revealing LLMs' struggles with multilingual code generation. Evaluation of 24 models uncovered Python overfitting and language-specific contamination.

June 20, 2026
TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement Technology

TERMS-Bench Diagnoses LLM Negotiation Agents Beyond Deal Rate for Enterprise Procurement

A new benchmark called TERMS-Bench goes beyond deal rate to diagnose why LLM negotiation agents fail, evaluating 13 frontier models on surplus extraction, cue use, belief calibration, and compliance. For enterprise procurement and trade, this offers actionable insights into AI agent weaknesses.

June 17, 2026
LLM-WikiRace Benchmark Reveals Frontier AI Models Still Struggle with Planning Over Knowledge Graphs Technology

LLM-WikiRace Benchmark Reveals Frontier AI Models Still Struggle with Planning Over Knowledge Graphs

Researchers introduced LLM-WikiRace, a benchmark to evaluate large language models on planning, reasoning, and world knowledge using Wikipedia hyperlinks. Top models like Gemini-3, GPT-5, and Claude Opus 4.5 achieve superhuman performance on easy tasks but drop sharply on hard difficulty, with Gemini-3 succeeding in only 23% of hard games. The study reveals that world knowledge helps only up to a point; beyond that, planning and long-horizon reasoning are the limiting factors.

June 16, 2026