iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline
Home ›› Technology ›› Ai ›› Robotics ›› FineVLA Framework Improves Robot Instruction Following by 62.7% in Real-World Dual-Arm Manipulation

FineVLA Framework Improves Robot Instruction Following by 62.7% in Real-World Dual-Arm Manipulation

Researchers introduce FineVLA, an open framework for fine-grained instruction alignment in vision-language-action (VLA) robot policies. The framework includes a dataset of 47,159 human-verified trajectories, a benchmark with 500 videos and 11,631 atomic facts, and a steerable policy that improves real-world dual-arm manipulation success from 49.9% (raw-only) to 62.7%.

iG
iGEN Editorial
June 16, 2026
FineVLA Framework Improves Robot Instruction Following by 62.7% in Real-World Dual-Arm Manipulation

Enterprise robotics deployments often struggle when robots must follow detailed execution instructions beyond simple goal-level commands. A new open framework called FineVLA, detailed in a paper on arXiv, addresses this gap by aligning robot actions with fine-grained human instructions about how tasks should be performed.

The framework, developed by researchers including Xintong Huang, Xuhong Zhang, Jinyu Yao, Yutong Sun, and Yuchong Wang, among others, targets a fundamental limitation in existing robot datasets: they typically pair trajectories with coarse goal-level language, leaving out execution-critical details such as active arm, approach direction, and contact region. This missing nuance limits steerable policy learning and robotic video understanding, according to the paper.

FineVLA Components and Dataset

FineVLA includes four main components: (1) a data construction tool that unifies 972,247 trajectories across 85,000 tasks from 10 open-source robot datasets; (2) a human-verified dataset called FineVLA-Data containing 47,159 fine-grained trajectories; (3) a held-out benchmark with 500 videos, 11,631 atomic facts, and 1,030 VQA questions; and (4) a robotics-specialized VLM annotator for scalable fine-grained annotation. The framework also includes a steerable VLA policy trained with controlled mixtures of fine-grained and raw goal-level instructions.

Experimental Results: Fine-Grained Supervision Boosts Success

The paper reports three key findings from experiments. First, fine-grained supervision does not sacrifice goal-level success: FineGrained-only outperforms Raw-only by +1.4 to +8.1 success-rate points across settings. Second, fine-grained and raw instructions are complementary, following a consistent inverted-U trend peaking at a FineGrained:Raw ratio of 1:2 to 1:1.

Setting FineGrained-Only Raw-Only Best Mixed (FG:Raw) Improvement
RoboTwin simulation baseline 86.8%/82.5% +?
Real-world dual-arm manipulation 49.9% 62.7% +12.8 points

In real-world dual-arm manipulation, the best mixed setting reached 62.7% success rate versus 49.9% for Raw-only, according to the paper. In RoboTwin simulation, the best mixed setting achieved 86.8% and 82.5% success rates.

Steerable Control Improvements

Third, fine-grained supervision improves steerable control. The largest real-world gains are observed on pose (+23 percentage points), color (+18), and approach direction (+18)—factors where goal-level instructions provide no guidance.

Overall, fine-grained language should augment goal-level instructions: specifying how to execute alongside what to achieve.

Implications for Enterprise Robotics

For enterprise technology decision-makers evaluating robotic automation, FineVLA demonstrates that incorporating detailed execution instructions can yield significant performance gains without complicating training. The open-source nature of the framework—including its dataset, benchmark, and policy infrastructure—allows organizations to test and adapt the approach for their own robotic systems, from warehouse picking to assembly line manipulation.

The researchers note that the framework uses controlled mixtures of fine-grained and raw instructions, with optimal results at a 1:2 to 1:1 ratio. This suggests that enterprises can augment existing goal-level command interfaces with more specific guidance to improve robot flexibility and task completion.

Future work could extend FineVLA to more complex industrial scenarios, though the paper does not detail specific integration paths or commercial availability. The framework is available at the project page linked in the paper.


Sources:

Keep Reading

Recommended Stories

New Training-Free Method Enables Robots to Follow Personalized Commands Like 'Bring My Cup' Technology

New Training-Free Method Enables Robots to Follow Personalized Commands Like 'Bring My Cup'

Researchers propose Visual Attentive Prompting (VAP), a training-free perceptual adapter that enables vision-language-action models to follow personalized commands by using reference images as visual prompts. VAP outperforms generic policies and token-learning baselines on simulation and real-world benchmarks.

July 8, 2026
New Benchmark and Method Address Occlusion in Vision-Language-Action Models for Robotics Technology

New Benchmark and Method Address Occlusion in Vision-Language-Action Models for Robotics

Researchers introduced LIBERO-Occ, an occlusion-oriented benchmark for Vision-Language-Action (VLA) models, and proposed Viewpoint Imagination (VIM), a method that generates a complementary view from an occluded primary observation to condition action prediction. Experiments show that state-of-the-art VLAs suffer substantial performance degradation under occlusion, and VIM improves robustness across task suites, occlusion types, and severity levels without requiring additional cameras at deployment.

June 16, 2026
For the First Time, Zoox Can Charge People for Rides in Its Steering-Wheel-Free Robotaxis Technology

For the First Time, Zoox Can Charge People for Rides in Its Steering-Wheel-Free Robotaxis

The National Highway Traffic Safety Administration (NHTSA) granted Amazon subsidiary Zoox a two-year exemption to deploy up to 5,000 steering-wheel-free robotaxis and charge for rides. Zoox will begin paid operations in Las Vegas, having already transported over 500,000 riders free of charge. The approval marks a milestone for autonomous vehicles built without traditional controls, subject to heightened oversight and safety reporting.

July 30, 2026
Google DeepMind's Gemini AI Now Controls Humanoid Robots for Dextrous Tasks Technology

Google DeepMind's Gemini AI Now Controls Humanoid Robots for Dextrous Tasks

Google DeepMind has released Gemini Robotics 2, an AI model that can control humanoid robots to perform complex physical tasks. The system combines vision language and action models, and has been demonstrated using Apptronik's Apollo 2 robot. Safety remains a key concern, with Google introducing a new benchmark called ASIMOV-Agentic.

July 30, 2026