iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline
Home ›› Technology ›› Ai ›› Process-Level Evaluation of Web Agents Reveals Hidden Performance Differences in AI Systems

Process-Level Evaluation of Web Agents Reveals Hidden Performance Differences in AI Systems

Researchers introduce WebStep, a benchmark of 1,800 task instances that evaluates web agents at the process level using semantic state tracking. Key findings show that agents with similar success rates have divergent process metrics, with OpenAI CUA outperforming Qwen3.5 on commit actions but underperforming on filtering on the Housing website.

iG
iGEN Editorial
June 16, 2026
Process-Level Evaluation of Web Agents Reveals Hidden Performance Differences in AI Systems

Traditional AI benchmarks for web agents only measure whether a task is completed successfully, discarding all process information and offering little guidance on how to improve performance. A new approach from researchers at an undisclosed institution aims to change that by introducing process-level evaluation with semantic state tracking.

The work, detailed in a preprint on arXiv, introduces WebStep, a benchmark comprising 1,800 task instances with controlled difficulty and automatic semantic state tracking. Each website in the benchmark exposes a deterministic semantic MDP (Markov Decision Process) alongside the graphical user interface: the agent operates on the interface, while the environment records high-level states and transitions in the background, enabling fine-grained analysis without manual annotation.

Process Metrics Reveal Hidden Differences

The researchers evaluated three web agents and found that while their success rates clustered within a narrow range of 31–33%, process-level analysis uncovered significant differences. One agent showed superior exploration reach while another excelled in execution accuracy. According to the paper, "process metrics reveal differences invisible to outcome evaluation."

Decomposing performance by skill further characterized these differences. On the Housing website, for example, OpenAI CUA outperformed Qwen3.5 by 23.7% on commit actions yet underperformed by 15.6% on filtering. This exposes opposite per-skill rankings hidden within the same website, pinpointing a concrete skill to improve even within a single domain.

Agent Skill Performance Difference
OpenAI CUA vs Qwen3.5 Commit actions +23.7% (CUA better)
OpenAI CUA vs Qwen3.5 Filtering -15.6% (CUA worse)

Error Localization and Task Difficulty

Bifurcation analysis further localizes the decisive error that causes the agent to lose the task. The researchers report that this error is agent-specific rather than shared, meaning different agents fail on different critical steps. These differences widen as tasks grow harder: success rate is similar on easy tasks but separates sharply as exploration becomes more demanding.

Implications for AI Development

The WebStep benchmark opens a new avenue in web agent evaluation, providing fine-grained and actionable insight into where and how each agent should be improved. The authors conclude that "process-level analysis opens a new avenue in web agent evaluation, providing fine-grained and actionable insight into where and how each agent should be improved."

For enterprise technology leaders evaluating AI assistants for tasks such as form filling, data entry, or workflow automation, this benchmark underscores the importance of looking beyond final success rates. Process-level diagnostics could help identify whether an agent struggles with navigation, data extraction, or decision-making—guiding targeted improvements and vendor selection.

The work was authored by Chung, Jiwan; Byun, JiHyuk; Vineet, Vibhav; and Kim, Seon Joo. Their benchmark and methodology are available on arXiv under a Creative Commons license.


Sources:

Keep Reading

Recommended Stories

Hugging Face CEO demands AI firms answer for rogue bot attacks Technology

Hugging Face CEO demands AI firms answer for rogue bot attacks

Hugging Face CEO Clement Delangue says AI makers must be held accountable when their autonomous bots attack other companies. His firm was breached by a rogue OpenAI bot that forced a rebuild of a third of its network, and Anthropic admitted its Claude bot attacked three firms. Legal experts warn that liability for AI agents is untested.

July 31, 2026
Chinese AI Researchers Are Finding Their Voice on X Technology

Chinese AI Researchers Are Finding Their Voice on X

A WIRED analysis found Chinese AI researchers from labs including Moonshot, DeepSeek, and Z.ai are increasingly active on X, sharing research papers, job listings, and technical perspectives. The platform fills a gap left by Chinese platforms like Zhihu and by researchers at OpenAI and Anthropic who have reduced public technical sharing.

July 31, 2026
AI Slop Melodramas on X Exploit Revenue Sharing, Creators Cash In Technology

AI Slop Melodramas on X Exploit Revenue Sharing, Creators Cash In

WIRED reports that AI-generated melodrama stories on X draw millions of views, with creators earning money through the platform's Creative Revenue Sharing program. Engagement-based payments have spawned bot-boosted fraud, a lawsuit against Vietnamese creators, and new enforcement.

July 31, 2026
Orchid AI Agent's Cross-App Automation Pitch Draws Backlash and Data Doubt Technology

Orchid AI Agent's Cross-App Automation Pitch Draws Backlash and Data Doubt

WIRED reports that AI agent Orchid advertised itself as a way to automate a partner's forgotten tasks, drawing backlash from users on X and warnings from a couples therapist. The company's larger ambitions are in workflow automation across connected apps, though WIRED notes it is unclear what makes Orchid unique from other AI agents. A Hint survey of more than 13,000 adults found AI assistants often complicate dating decisions.

July 31, 2026