iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› Robotics ›› ScoutVLA: New Dual-Expert AI Model Boosts UAV Active Perception for Embodied Question Answering

ScoutVLA: New Dual-Expert AI Model Boosts UAV Active Perception for Embodied Question Answering

Researchers introduce ScoutVLA, a vision-language-action model for UAV active perception, achieving 10.48x higher strict success rate and 7.72x higher QA correctness over baselines. The model features a decoupled dual-expert architecture inspired by scout bee waggle dance.

iG
iGEN Editorial
June 16, 2026
ScoutVLA: New Dual-Expert AI Model Boosts UAV Active Perception for Embodied Question Answering

Unmanned Aerial Vehicles (UAVs) are increasingly used in inspection and surveillance, but their ability to answer natural language questions by actively exploring environments – a task known as Embodied Question Answering (EQA) – has been limited. Existing outdoor EQA systems typically stop once a target enters the UAV's field of view, leaving fine-grained viewpoint adjustments for evidence-seeking questions unresolved. According to a new preprint on arXiv, researchers have introduced ScoutVLA, an evidence-driven Vision-Language-Action (VLA) model designed to address this gap.

"To address this issue, we introduce FG-EQA, a fine-grained active perception EQA benchmark with more than 40K simulated trajectories and 1K real-world trajectories." – from the paper

FG-EQA Benchmark: Fine-Grained Active Perception

The team first developed FG-EQA, a benchmark specifically for fine-grained active perception. It includes over 40,000 simulated trajectories and 1,000 real-world trajectories, providing a robust testing ground for EQA systems. The benchmark challenges UAVs to not only locate targets but also refine their viewpoint to gather evidence for answering questions.

ScoutVLA Architecture: Dual-Expert Design

ScoutVLA draws inspiration from the "waggle dance" of scout bees, which iteratively adjust flight paths to verify target information. The model employs a decoupled dual-expert architecture:

  • A vision-language expert that infers semantic intent to identify missing evidence.
  • An independent action expert that uses high-DoF flow matching to generate continuous viewpoint-refinement trajectories.

This separation is key to balancing the competing demands of continuous control and semantic reasoning.

Training and Results

To avoid interference between the two experts, the researchers devised a decoupled training strategy with a knowledge insulation mechanism that prevents action gradients from erasing the model's multimodal reasoning ability. The results show significant improvements over state-of-the-art baselines:

Metric Improvement over Baselines
Average strict success rate 10.48× higher
Average QA correctness 7.72× higher

These gains were demonstrated in extensive simulated experiments.

Real-World Validation

A qualitative real-world field study also confirmed ScoutVLA's superiority. While the paper does not provide specific field results, it states the study "verifies the superiority of ScoutVLA over the state-of-the-art baselines."

Implications for Autonomous Systems

ScoutVLA represents a step forward in UAV active perception, with potential applications in infrastructure inspection, search and rescue, and precision agriculture. For enterprise technology leaders, the model's ability to reason about missing evidence and adjust viewpoints autonomously could reduce the need for manual drone piloting and enable more sophisticated autonomous missions. The decoupled architecture also offers a blueprint for integrating vision-language reasoning with continuous control in other robotic domains.


Sources:

Keep Reading

Recommended Stories

New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models Technology

New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models

A new computer vision paper from arXiv investigates how visual tokens are integrated into large language models (LLMs) under two paradigms: in-context prompting and layer-wise injection. The authors find that visual tokens enter the LLM as 'disguised visual context' lacking linguistic structure, then evolve differently depending on the integration architecture. They show that attention allocation alone is insufficient, and performance depends on the quality of visual representations at each layer.

July 8, 2026
STAR Allocation Method Improves Text-to-Image AI Training with Spatiotemporal Rewards Technology

STAR Allocation Method Improves Text-to-Image AI Training with Spatiotemporal Rewards

A new method called SpatioTemporal Adaptive Reward (STAR) Allocation improves reinforcement learning post-training for text-to-image generation. By using text-image attention to allocate rewards to relevant latent regions, STAR enhances compositional semantic alignment, text rendering, and preference optimization without changing the external reward source. The method was validated on Stable Diffusion 3.5 Medium, achieving top scores on GenEval, OCR, and PickScore benchmarks.

June 20, 2026
RoboSSM Introduces State-Space Models for Scalable In-Context Imitation Learning in Robotics Technology

RoboSSM Introduces State-Space Models for Scalable In-Context Imitation Learning in Robotics

RoboSSM is a new method for in-context imitation learning (ICIL) that replaces Transformer-based architectures with state-space models (SSMs). The approach uses Longhorn, a state-of-the-art SSM, enabling linear-time inference and strong extrapolation to longer prompts. Experiments on the LIBERO benchmark show improved generalization to unseen and long-horizon tasks compared to Transformer-based ICIL methods.

June 20, 2026
Lagrange: New Open-Vocabulary Sparse Framework Promises Robust Autonomous Driving in Open Worlds Technology

Lagrange: New Open-Vocabulary Sparse Framework Promises Robust Autonomous Driving in Open Worlds

A new framework called Lagrange, based on Masked Latent Fields and vision-language models, aims to enable autonomous vehicles to handle out-of-distribution scenarios and produce kinematically valid trajectories. Offline evaluations on nuScenes and CODA benchmarks show promising results for robust open-world driving.

June 20, 2026