iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› Computer Vision ›› New Benchmark and Method Address Occlusion in Vision-Language-Action Models for Robotics

New Benchmark and Method Address Occlusion in Vision-Language-Action Models for Robotics

Researchers introduced LIBERO-Occ, an occlusion-oriented benchmark for Vision-Language-Action (VLA) models, and proposed Viewpoint Imagination (VIM), a method that generates a complementary view from an occluded primary observation to condition action prediction. Experiments show that state-of-the-art VLAs suffer substantial performance degradation under occlusion, and VIM improves robustness across task suites, occlusion types, and severity levels without requiring additional cameras at deployment.

iG
iGEN Editorial
June 16, 2026
New Benchmark and Method Address Occlusion in Vision-Language-Action Models for Robotics

Vision-Language-Action (VLA) models have achieved strong performance on standard manipulation benchmarks, but most evaluations assume that task-relevant objects are fully visible. According to the paper "LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination," this assumption often fails in realistic settings, where occlusion makes manipulation partially observable. The authors introduced LIBERO-Occ, an occlusion-oriented extension of the LIBERO benchmark, and Viewpoint Imagination (VIM), a method that generates a complementary view from an occluded primary observation and conditions action prediction on both observed and imagined evidence. Experiments show that state-of-the-art VLAs suffer substantial performance degradation under occlusion, and VIM improves robustness across task suites, occlusion types, and severity levels without requiring additional cameras at deployment time.

The Occlusion Challenge for Vision-Language-Action Models

VLA models integrate visual perception, language understanding, and action generation for robotic manipulation. Standard benchmarks typically present scenes where task-relevant objects are fully visible, a condition that rarely holds in real-world deployments. The paper identifies scene-induced occlusion as a fundamental challenge for VLA models. In settings such as cluttered bins, shelves, or industrial environments, objects may be partially hidden by other items or by the robot's own gripper. The authors report that state-of-the-art VLAs experience substantial performance degradation when occlusion is present, underscoring the need for robust perception-completion mechanisms.

LIBERO-Occ: A Benchmark for Scene-Induced Occlusion

To systematically evaluate VLA models under occlusion, the researchers created LIBERO-Occ, an occlusion-oriented extension of the existing LIBERO benchmark. This new benchmark introduces various occlusion types and severity levels across multiple manipulation task suites. The paper states that LIBERO-Occ is designed to assess how well VLAs handle partially observable conditions. The benchmark and corresponding code are publicly available, enabling the research community to test and compare occlusion-robust methods.

Viewpoint Imagination (VIM): Technical Overview

The proposed method, Viewpoint Imagination (VIM), addresses occlusion by generating a complementary view from the primary occluded observation. VIM conditions action prediction on both the observed and the imagined evidence, effectively providing the model with a more complete scene understanding. According to the authors, this approach improves robustness across task suites, occlusion types, and severity levels. Importantly, VIM does not require additional cameras at deployment time, meaning it can be applied to existing robotic systems without hardware modifications. The paper suggests that viewpoint imagination is a promising mechanism for perception completion in partially observable manipulation.

Implications for Robotics in Logistics and Supply Chain

Although the experiments are conducted on manipulation benchmarks, the principles of LIBERO-Occ and VIM are directly relevant to robotics in logistics and supply chain environments. Occlusion is a common occurrence in warehouse automation, such as when a robotic arm picks items from cluttered bins or when a mobile robot navigates tightly packed shelves. The ability to generate a complementary view without extra cameras could improve the reliability of automated picking, packing, and sorting operations. The research provides a foundation for developing VLA models that are more resilient to the imperfect viewing conditions typical of industrial settings.

Component Description
LIBERO-Occ Occlusion-oriented extension of LIBERO benchmark for evaluating VLA models under scene-induced occlusion
Viewpoint Imagination (VIM) Generates complementary view from occluded primary observation; conditions action prediction on observed and imagined evidence
Key Result State-of-the-art VLAs suffer substantial degradation under occlusion; VIM improves robustness without additional cameras

The paper's authors — Li, Taishan; Zhang, Jiwen; Wang, Siyuan; Huang, Xuanjing; and Wei, Zhongyu — have released the benchmark and code at this link.


Sources:

Keep Reading

Recommended Stories

New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models Technology

New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models

A new computer vision paper from arXiv investigates how visual tokens are integrated into large language models (LLMs) under two paradigms: in-context prompting and layer-wise injection. The authors find that visual tokens enter the LLM as 'disguised visual context' lacking linguistic structure, then evolve differently depending on the integration architecture. They show that attention allocation alone is insufficient, and performance depends on the quality of visual representations at each layer.

July 8, 2026
New Training-Free Method Enables Robots to Follow Personalized Commands Like 'Bring My Cup' Technology

New Training-Free Method Enables Robots to Follow Personalized Commands Like 'Bring My Cup'

Researchers propose Visual Attentive Prompting (VAP), a training-free perceptual adapter that enables vision-language-action models to follow personalized commands by using reference images as visual prompts. VAP outperforms generic policies and token-learning baselines on simulation and real-world benchmarks.

July 8, 2026
New AI Research Shows Vision-Language Models Think Better with Visual Grounding Technology

New AI Research Shows Vision-Language Models Think Better with Visual Grounding

Researchers introduce visually grounded thinking, a reasoning process that interleaves natural-language thoughts with explicit point or box groundings to image regions. The method, using a scalable synthesis pipeline and grounding-aware reinforcement learning, consistently improves performance of Gemma3-4B-IT on counting and spatial reasoning benchmarks, with the 4B model matching or surpassing the 27B variant.

June 21, 2026
DF3DV-1K: Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis Technology

DF3DV-1K: Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Researchers introduced DF3DV-1K, a large-scale real-world dataset with 1,048 scenes and 89,924 images for distractor-free novel view synthesis. The dataset spans 128 distractor types and 161 scene themes, enabling benchmarking of nine radiance field methods and 3D Gaussian Splatting. Fine-tuning a diffusion-based 2D enhancer on DF3DV-1K achieved average improvements of 0.96 dB PSNR and 0.057 LPIPS.

June 21, 2026