Topic
action
ROSE Benchmark Reveals Perception-to-Action Gap in Multimodal AI Models
The ROSE benchmark measures how reliably multimodal large language models (MLLMs) convert visual evidence into context-appropriate actions. Testing nine recent models, researchers found performance drops of up to 44.5 percentage points from counting to region-conditioned action, while humans achieve 98.8% accuracy.
New Robotic Architecture AVP Improves Pick-and-Place Success Rate by 37% over Existing Models
A new research paper introduces AVP (Action with Visual Primitives), an end-to-end architecture for robotic manipulation that decouples visual-language reasoning from action generation. In real-robot pick-and-place experiments, AVP achieved a 37.04% higher success rate than the pi_0.5 baseline, with gains in data efficiency, spatial-compositional generalization, and object-level transfer.
Survey on Medical Embodied AI Highlights Integration of Perception, Decision-Making, and Action
A systematic survey of medical embodied AI examines its core components — perception, decision-making, and action — and their coordinated integration for real-world clinical workflows. The paper reviews representative applications, datasets, and challenges, highlighting the need for unified system-level organization beyond individual functional aspects.