iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face Inside the rogue ChatGPT hack of Hugging Face: AI agents operate at superhuman speed but make clumsy mistakes Landstar Expects to Emerge a Winner After Supreme Court’s Montgomery Ruling Widens Broker Liability New Senate bill targets 'chameleon carriers' that reopen to escape penalties Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face Inside the rogue ChatGPT hack of Hugging Face: AI agents operate at superhuman speed but make clumsy mistakes Landstar Expects to Emerge a Winner After Supreme Court’s Montgomery Ruling Widens Broker Liability New Senate bill targets 'chameleon carriers' that reopen to escape penalties Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb
Home ›› Technology ›› Ai ›› Llms ›› ROSE Benchmark Reveals Perception-to-Action Gap in Multimodal AI Models

ROSE Benchmark Reveals Perception-to-Action Gap in Multimodal AI Models

The ROSE benchmark measures how reliably multimodal large language models (MLLMs) convert visual evidence into context-appropriate actions. Testing nine recent models, researchers found performance drops of up to 44.5 percentage points from counting to region-conditioned action, while humans achieve 98.8% accuracy.

iG
iGEN Editorial
June 22, 2026
ROSE Benchmark Reveals Perception-to-Action Gap in Multimodal AI Models

Multimodal large language models (MLLMs) are increasingly deployed to act on visual information, but a new benchmark reveals a significant gap between perception and context-specific action. The ROSE (Reference-conditioned Oddity and Symbolic Execution) benchmark, described in a paper on arXiv, systematically tests how reliably a model can turn the same visual evidence into the action required by the current task context.

The Perception-to-Action Gap

According to the paper, ROSE holds the visual scene fixed while varying region constraints and required symbolic outputs. Through coupled counting and coordinate-action tasks, the benchmark tests whether models can infer an implicit majority reference and act on resulting fine-grained visual evidence under changing contexts. Across nine recent MLLMs, performance drops by as much as 44.5 percentage points from counting-oriented tasks to region-conditioned action. In contrast, human performance stands at 98.8%. The gap persists even on paired scenes and regions for which the same model returns the correct count.

How ROSE Works

The benchmark uses a controlled setup: the visual scene remains constant, but the task switches between counting objects and performing actions based on coordinates. This isolates the model's ability to adapt shared visual evidence to different task contexts. Global-click and matched local controls show that coordinate grounding explains only part of the loss, revealing a distinct, model-dependent bottleneck in turning shared visual evidence into context-specific actions.

Implications for Enterprise AI

For enterprise technology leaders evaluating MLLMs for automation tasks — such as visual inspection in logistics or document processing in trade — the ROSE findings highlight a critical reliability issue. A model that correctly counts items in a scene may fail to act on that same scene when required to, for example, select a specific object or region. The 44.5 percentage point drop underscores that current MLLMs lack robust context-sensitive decision-making, a prerequisite for deployment in safety-critical or high-stakes supply chain operations.

Model-Level Variability

The paper tested nine recent MLLMs, but does not name specific models in the abstract. The model-dependent bottleneck suggests that different architectures or training regimes handle the perception-to-action transition differently. Enterprises should demand benchmarks like ROSE that isolate this specific capability, rather than relying solely on general vision-language accuracy metrics.

"How reliably can a model turn the same visual evidence into the action required by the current context?" — This question, posed by the authors, is central to the ROSE benchmark and to enterprise AI deployment.

As MLLMs move into production environments, benchmarks that expose the perception-to-action gap become essential procurement criteria. ROSE provides a controlled, reproducible method for assessing this capability, contributing to more trustworthy AI for trade, logistics, and beyond.


Sources:

Keep Reading

Recommended Stories

SLUM-i: AI Semi-Supervised Learning Maps Informal Settlements with Benchmark Dataset Technology

SLUM-i: AI Semi-Supervised Learning Maps Informal Settlements with Benchmark Dataset

A new AI framework called SLUM-i uses semi-supervised learning to map informal settlements in cities like Lahore, Karachi, and Mumbai. It introduces a benchmark dataset and achieves up to +5.9 pp mIoU improvement over existing methods.

June 17, 2026
Deep Residual Injection Method Enables Full-Spectrum Forensic AI Detection in Multimodal Models Technology

Deep Residual Injection Method Enables Full-Spectrum Forensic AI Detection in Multimodal Models

Researchers propose Deep Visual Residual MLLM (Deep-VRM), a method that injects low-level artifact signals into multimodal large language models without disrupting pre-trained semantic knowledge. The approach achieves state-of-the-art detection of AI-generated images across multiple benchmarks.

June 16, 2026
Research Shows 'Retrieve, Don't Retrain' Approach Cuts AI Model Adaptation Costs Technology

Research Shows 'Retrieve, Don't Retrain' Approach Cuts AI Model Adaptation Costs

A new research paper from arXiv proposes a retrieval-augmented vision-language-action (VLA) policy that eliminates the need for per-task fine-tuning. By retrieving relevant demonstrations from a pool at test time, the frozen policy adapts to new tasks without updating model parameters. The method shows strong results on robotic manipulation benchmarks, including PushT and RoboTwin 2.0, and on a real robot.

June 16, 2026
UrbanWell Benchmark Puts Multimodal LLMs to Test on Spatio-Temporal Urban Wellbeing Analytics Technology

UrbanWell Benchmark Puts Multimodal LLMs to Test on Spatio-Temporal Urban Wellbeing Analytics

Researchers introduce UrbanWell, a large-scale benchmark for evaluating multimodal large language models on spatio-temporal urban wellbeing analytics. The benchmark covers 38 cities, multiple years, and diverse indicators including environment, accessibility, urban form, vitality, and subjective perception. Testing 15 state-of-the-art MLLMs in zero-shot settings reveals substantial performance variations across heterogeneous indicators.

June 16, 2026