iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Topics ›› ai evaluation

Topic

ai evaluation

5 stories
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows Technology
Artificial Intelligence #voice agents#post-interruption recovery

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

IHBench, a new benchmark from researchers including Salimi et al., evaluates how voice agents recover after interruptions in structured enterprise workflows. The benchmark tests 27 audio-language models from OpenAI, Google, and the open-weight community, finding that closed-weight models are consistently more robust, degrading 3.3x more slowly in long conversations.

Jun 22, 2026 1 source
RTSGameBench Benchmark Tests Strategic Reasoning in Vision-Language Models Technology
Artificial Intelligence #rtsgamebench#benchmark

RTSGameBench Benchmark Tests Strategic Reasoning in Vision-Language Models

A new benchmark called RTSGameBench evaluates strategic reasoning in vision-language models (VLMs) using the real-time strategy game Beyond All Reason. The benchmark includes diagnostic mini-games, diverse matchup structures, and a self-evolving generation framework. Initial tests show state-of-the-art VLMs struggle with tighter coordination, multiagent tasks, and increased scale.

Jun 21, 2026 1 source
The Scaffold Effect: How Prompt Framing Skews AI Evaluation in Clinical Vision-Language Models Technology
Artificial Intelligence #artificial intelligence#vision-language models

The Scaffold Effect: How Prompt Framing Skews AI Evaluation in Clinical Vision-Language Models

A study on arXiv evaluating 12 open-weight vision-language models (VLMs) on clinical neuroimaging datasets found that up to 58% of apparent multimodal performance gains are due to prompt framing rather than genuine reasoning. The researchers identified a 'scaffold effect' where merely mentioning MRI availability in the task prompt accounts for 70-80% of F1 improvement, even when no imaging data is present. Expert evaluation also revealed fabrication of neuroimaging-grounded justifications, raising concerns about the reliability of VLM evaluations in clinical settings.

Jun 20, 2026 1 source
CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models Technology
Artificial Intelligence #large language models#combinatorial counting

CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models

CombEval is a dynamic benchmark for evaluating combinatorial counting in large language models. It uses typed Cofola specifications to generate problems with verified answers. Tests on 11 LLMs reveal persistent failures on ordered objects, indistinguishable elements, and nested dependencies.

Jun 20, 2026 1 source
Process-Level Evaluation of Web Agents Reveals Hidden Performance Differences in AI Systems Technology
Artificial Intelligence #web agents#process-level evaluation

Process-Level Evaluation of Web Agents Reveals Hidden Performance Differences in AI Systems

Researchers introduce WebStep, a benchmark of 1,800 task instances that evaluates web agents at the process level using semantic state tracking. Key findings show that agents with similar success rates have divergent process metrics, with OpenAI CUA outperforming Qwen3.5 on commit actions but underperforming on filtering on the Housing website.

Jun 16, 2026 1 source