iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Maharashtra’s ₹500 crore AI agriculture policy targets data, traceability and farm advisory Commercial LPG prices drop: 19-kg cylinder rate cut by ₹202 in Delhi, ₹209 in Kolkata Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Maharashtra’s ₹500 crore AI agriculture policy targets data, traceability and farm advisory Commercial LPG prices drop: 19-kg cylinder rate cut by ₹202 in Delhi, ₹209 in Kolkata Commercial LPG Prices Cut by Over Rs 200; Delhi, Kolkata 19-kg Cylinder Rates Published US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue
Home ›› Technology ›› Ai ›› Computer Vision ›› New Robotic Architecture AVP Improves Pick-and-Place Success Rate by 37% over Existing Models

New Robotic Architecture AVP Improves Pick-and-Place Success Rate by 37% over Existing Models

A new research paper introduces AVP (Action with Visual Primitives), an end-to-end architecture for robotic manipulation that decouples visual-language reasoning from action generation. In real-robot pick-and-place experiments, AVP achieved a 37.04% higher success rate than the pi_0.5 baseline, with gains in data efficiency, spatial-compositional generalization, and object-level transfer.

iG
iGEN Editorial
June 17, 2026
New Robotic Architecture AVP Improves Pick-and-Place Success Rate by 37% over Existing Models

Robotic manipulation systems that integrate vision, language, and action often force a single model to handle reasoning and motor control simultaneously, limiting learning efficiency and generalization. A new architecture called AVP (Action with Visual Primitives) addresses this by separating high-level task understanding from low-level action execution.

According to the research paper published on arXiv, AVP uses a vision-language model (VLM) to infer the next-stage target and emit visual-primitive tokens. These tokens condition a flow-matching action expert, which generates precise end-effector movements. Supervision is derived from end-effector kinematics, ensuring the action expert learns motor control directly without relearning cognitive capabilities already present in the pretrained VLM.

Key Results from Real-Robot Experiments

The researchers conducted real-robot experiments on general pick-and-place tasks. AVP improved the success rate by 37.04% over the pi_0.5 baseline, and also outperformed other recent methods. The experiments demonstrated consistent gains in several areas:

  • Data efficiency: AVP required fewer demonstrations to achieve comparable performance.
  • Spatial-compositional generalization: The robot could handle novel spatial arrangements of objects.
  • Object-level transfer: Skills learned on one object transferred to different objects without retraining.
Metric AVP vs pi_0.5
Success rate improvement +37.04%
Data efficiency Higher (fewer demos)
Spatial-compositional generalization Superior
Object-level transfer Superior

How AVP Works

AVP's design decouples the typical VLA (Vision-Language-Action) pipeline. The VLM processes the language instruction and visual observations to generate visual-primitive tokens, which act as an intermediate representation. The action expert then uses these tokens to produce continuous motor commands via flow matching. This separation prevents the action expert from having to implicitly relearn perception and reasoning, which are already handled by the VLM.

Implications for Enterprise Robotics

For technology leaders in logistics and supply chain, AVP's approach could lead to more adaptive and data-efficient warehouse automation. The ability to generalize to new objects and spatial configurations reduces the need for extensive retraining per SKU. The 37% improvement in pick-and-place success translates directly to fewer interventions and higher throughput in fulfillment centers.

Future Outlook

While AVP has been demonstrated only on pick-and-place tasks, the underlying architecture is extensible to other manipulation problems. The paper's authors—Guo, Weilong; Wang, Yuchen; Zhou, Renping; Zhang, Yunfeng; Fang, Rui; Pang, Yuyang; Xu, Wenda; and Huang, Gao—did not specify commercial partnerships. However, the open availability of the paper on arXiv suggests potential for adoption in research and eventually in commercial robotics platforms.

By separating visual reasoning from motor control, AVP addresses a fundamental bottleneck in VLA models. For enterprises investing in robotic automation, this method offers a pathway to more robust and generalizable systems that require less task-specific data.


Sources:

Keep Reading

Recommended Stories

SARLO-80: New Dataset Combines Very-High-Resolution SAR and Optical Imagery with Language Descriptions Technology

SARLO-80: New Dataset Combines Very-High-Resolution SAR and Optical Imagery with Language Descriptions

Researchers have released SARLO-80, a large-scale dataset combining very-high-resolution synthetic aperture radar (SAR) imagery, aligned optical imagery, and natural-language descriptions. Built from Umbra spotlight acquisitions, the dataset contains 119,566 triplets across 72 countries, standardized to 80cm slant-range resolution. It aims to advance multimodal foundation models for SAR by providing complex-valued measurements and native acquisition geometry.

July 8, 2026
New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models Technology

New Research Reveals How Visual Tokens Evolve Inside Vision-Language Models

A new computer vision paper from arXiv investigates how visual tokens are integrated into large language models (LLMs) under two paradigms: in-context prompting and layer-wise injection. The authors find that visual tokens enter the LLM as 'disguised visual context' lacking linguistic structure, then evolve differently depending on the integration architecture. They show that attention allocation alone is insufficient, and performance depends on the quality of visual representations at each layer.

July 8, 2026
Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis Technology

Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis

Researchers propose an object-centric OOD detection framework that leverages object co-occurrence patterns to overcome simplicity bias, achieving competitive results on near-OOD and full-spectrum settings.

July 8, 2026
New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs Technology

New Framework GeoVR Learns 3D Spatial Intelligence from 2D Videos for Multimodal LLMs

Multimodal Large Language Models (MLLMs) traditionally lack intrinsic 3D awareness. Researchers present GeoVR, a framework that learns geometric representations from 2D video sequences, restructuring the semantic latent space to unlock spatial intelligence. GeoVR uses four complementary geometric targets from pre-trained 3D foundation models, achieving state-of-the-art performance on spatial reasoning benchmarks.

July 8, 2026