iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition Relay Q: London Startup's AI Microphone Puts Hands-Free Voice Dictation on the Desktop Google Pixel 10a Crowned Best Budget Pixel in WIRED's Updated 2026 Buying Guide Global Steel Wire seeks fresh Santander terminal concession Veritas Shipmanagement books fresh ultramax pair at COSCO yard, Splash247 reports Seanergy linked to fresh newcastlemax at Hengli as dry bulk orderbook grows Weaker rupee may push foreign assets over FAST-DS Rs 1 crore limit, raising tax bill 45 Indian power plants face critically low coal stocks as monsoon hits supply SFL Makes Fresh $363m Car Carrier Play With Four LNG Dual-Fuel Newbuilds Iran Blacklist Threatens Hormuz Shuttle Tanker Lifeline for Gulf Crude Keyfield International Enters Dredging Market with $24.7m Vessel Acquisition
Home ›› Technology ›› Ai ›› Computer Vision ›› New Robotic Architecture AVP Improves Pick-and-Place Success Rate by 37% over Existing Models

New Robotic Architecture AVP Improves Pick-and-Place Success Rate by 37% over Existing Models

A new research paper introduces AVP (Action with Visual Primitives), an end-to-end architecture for robotic manipulation that decouples visual-language reasoning from action generation. In real-robot pick-and-place experiments, AVP achieved a 37.04% higher success rate than the pi_0.5 baseline, with gains in data efficiency, spatial-compositional generalization, and object-level transfer.

iG
iGEN Editorial
June 17, 2026
New Robotic Architecture AVP Improves Pick-and-Place Success Rate by 37% over Existing Models

Robotic manipulation systems that integrate vision, language, and action often force a single model to handle reasoning and motor control simultaneously, limiting learning efficiency and generalization. A new architecture called AVP (Action with Visual Primitives) addresses this by separating high-level task understanding from low-level action execution.

According to the research paper published on arXiv, AVP uses a vision-language model (VLM) to infer the next-stage target and emit visual-primitive tokens. These tokens condition a flow-matching action expert, which generates precise end-effector movements. Supervision is derived from end-effector kinematics, ensuring the action expert learns motor control directly without relearning cognitive capabilities already present in the pretrained VLM.

Key Results from Real-Robot Experiments

The researchers conducted real-robot experiments on general pick-and-place tasks. AVP improved the success rate by 37.04% over the pi_0.5 baseline, and also outperformed other recent methods. The experiments demonstrated consistent gains in several areas:

  • Data efficiency: AVP required fewer demonstrations to achieve comparable performance.
  • Spatial-compositional generalization: The robot could handle novel spatial arrangements of objects.
  • Object-level transfer: Skills learned on one object transferred to different objects without retraining.
Metric AVP vs pi_0.5
Success rate improvement +37.04%
Data efficiency Higher (fewer demos)
Spatial-compositional generalization Superior
Object-level transfer Superior

How AVP Works

AVP's design decouples the typical VLA (Vision-Language-Action) pipeline. The VLM processes the language instruction and visual observations to generate visual-primitive tokens, which act as an intermediate representation. The action expert then uses these tokens to produce continuous motor commands via flow matching. This separation prevents the action expert from having to implicitly relearn perception and reasoning, which are already handled by the VLM.

Implications for Enterprise Robotics

For technology leaders in logistics and supply chain, AVP's approach could lead to more adaptive and data-efficient warehouse automation. The ability to generalize to new objects and spatial configurations reduces the need for extensive retraining per SKU. The 37% improvement in pick-and-place success translates directly to fewer interventions and higher throughput in fulfillment centers.

Future Outlook

While AVP has been demonstrated only on pick-and-place tasks, the underlying architecture is extensible to other manipulation problems. The paper's authors—Guo, Weilong; Wang, Yuchen; Zhou, Renping; Zhang, Yunfeng; Fang, Rui; Pang, Yuyang; Xu, Wenda; and Huang, Gao—did not specify commercial partnerships. However, the open availability of the paper on arXiv suggests potential for adoption in research and eventually in commercial robotics platforms.

By separating visual reasoning from motor control, AVP addresses a fundamental bottleneck in VLA models. For enterprises investing in robotic automation, this method offers a pathway to more robust and generalizable systems that require less task-specific data.


Sources:

Keep Reading

Recommended Stories

SARLO-80: New Dataset Combines Very-High-Resolution SAR and Optical Imagery with Language Descriptions Technology

SARLO-80: New Dataset Combines Very-High-Resolution SAR and Optical Imagery with Language Descriptions

Researchers have released SARLO-80, a large-scale dataset combining very-high-resolution synthetic aperture radar (SAR) imagery, aligned optical imagery, and natural-language descriptions. Built from Umbra spotlight acquisitions, the dataset contains 119,566 triplets across 72 countries, standardized to 80cm slant-range resolution. It aims to advance multimodal foundation models for SAR by providing complex-valued measurements and native acquisition geometry.

July 8, 2026
New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models Technology

New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models

A new computer vision paper from arXiv investigates how visual tokens are integrated into large language models (LLMs) under two paradigms: in-context prompting and layer-wise injection. The authors find that visual tokens enter the LLM as 'disguised visual context' lacking linguistic structure, then evolve differently depending on the integration architecture. They show that attention allocation alone is insufficient, and performance depends on the quality of visual representations at each layer.

July 8, 2026
Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis Technology

Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis

Researchers propose an object-centric OOD detection framework that leverages object co-occurrence patterns to overcome simplicity bias, achieving competitive results on near-OOD and full-spectrum settings.

July 8, 2026
New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs Technology

New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs

Multimodal Large Language Models (MLLMs) traditionally lack intrinsic 3D awareness. Researchers present GeoVR, a framework that learns geometric representations from 2D video sequences, restructuring the semantic latent space to unlock spatial intelligence. GeoVR uses four complementary geometric targets from pre-trained 3D foundation models, achieving state-of-the-art performance on spatial reasoning benchmarks.

July 8, 2026