iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
New Senate bill targets 'chameleon carriers' that reopen to escape penalties Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity New Senate bill targets 'chameleon carriers' that reopen to escape penalties Werner Enterprises Posts Highest Revenue Per Truck Growth in One-Way Segment in a Decade CMA CGM and Stonepeak Launch United Ports LLC in $2.4 Billion Terminal Joint Venture UPS shift away from Amazon shows bigger payoff Lanesurf: 62% of Loads Get Vetted Carrier Offers Before Brokers Arrive India-China Border Trade Via Lipulekh Resumes Aug 1; China Permits 20 Traders Geopolitics Drives CMA CGM Q2 Profit Surge of 42% as Volumes and Rates Climb Benchmark Diesel Price Rises Third Week as Futures Plunge; Spread Hits Record Indian Government Limits Sugar Dealers to 400 Tonnes Stock Until November to Curb Hoarding Tenants signing longer leases for larger warehouses as 3PLs lock in capacity
Home ›› Technology ›› Ai ›› Robotics ›› New Training-Free Method Compresses Vision-Language-Action Models by 50% Without Performance Loss

New Training-Free Method Compresses Vision-Language-Action Models by 50% Without Performance Loss

A research team led by Gia-Binh Ho et al. discovered that Vision-Language-Action (VLA) models exhibit severe layer-wise redundancy. They introduced a training-free compression pipeline using Centered Kernel Alignment to remove twin layers, achieving up to 50% depth reduction, 40-50% faster fine-tuning, and 30% faster inference while matching or exceeding full-scale performance.

iG
iGEN Editorial
June 20, 2026
New Training-Free Method Compresses Vision-Language-Action Models by 50% Without Performance Loss

Vision-Language-Action (VLA) models, which combine vision, language, and action capabilities, have become foundational for robotic manipulation. However, their multi-billion parameter architectures demand massive computational resources for fine-tuning and real-time inference. A new study from researchers at multiple institutions reveals that these models contain significant layer-wise redundancy, enabling a training-free compression that reduces model depth by up to 50% while maintaining or improving performance.

The paper, titled "Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think" and authored by Gia-Binh Ho, Trong-Bao Ha, Thien-Loc Vo, Khoa Møller, Philip Lund, Quang T Dinh, Long Dam, Tuan Duong, Vu Luu, Tung M Le, Trung Tran Nguyen, Minh An Thai, Ngan Sonntag, Daniel Zou, James Peters, Jan Duy M H, and Vien Ngo Anh, introduces a structural compression pipeline that is entirely training-free. According to the paper, existing methods require loading full-scale models to learn optimized token reductions or dynamic layer selectors, but the new approach bypasses this need entirely.

A Single Forward Pass to Identify Redundancy

The core innovation is the use of Centered Kernel Alignment (CKA) to identify redundant layer features with just a single forward pass. The researchers found that models like pi_0 and GR00T-N1.5 exhibit severe layer-wise representational redundancy despite being trained on diverse physical trajectories. By removing twin layers, the pipeline permanently compresses model depth by up to 50% across both the VLM backbone and the continuous control policy head.

Dual Acceleration Benefits

Downstream fine-tuning of the streamlined architecture yields a dual acceleration benefit: a 40-50% reduction in training time and up to 30% faster real-time inference, according to the study. Importantly, the compressed model matches or exceeds the performance of the full-scale base model, challenging the assumption that more layers are always better for VLA models.

Validation Across Diverse Benchmarks

The method was comprehensively validated across three simulation benchmarks (LIBERO, RoboCasa, SimplerEnv) and 10 diverse real-world manipulation tasks across 4 unique robotic embodiments. The results consistently showed that advanced VLAs require significantly fewer layers than previously assumed, offering a highly compute-efficient paradigm for scalable robot learning.

Implications for Deployment

The training-free nature of the compression is particularly valuable for enterprise settings where computational budgets are constrained. By eliminating the need to load full-scale models for optimization, the pipeline reduces infrastructure costs and simplifies deployment of VLA models in real-time systems. For logistics and supply chain applications involving robotic manipulation, this could accelerate the adoption of AI-driven automation without additional hardware investment.

The paper is available on arXiv under the identifier 2606.20246. The authors have not announced plans to release the code or pre-compressed models, but the method is described in sufficient detail for other teams to replicate the results.


Sources:

Keep Reading

Recommended Stories

REST-GAN: A Deep Generative Model for Resting-State EEG Synthesis and Transferable Representation Learning Technology

REST-GAN: A Deep Generative Model for Resting-State EEG Synthesis and Transferable Representation Learning

Researchers introduce REST-GAN, a generative adversarial network for resting-state EEG that both synthesizes realistic neural signals and learns transferable representations. The model achieves high precision and recall in band-power features and shows competitive performance in demographic classification tasks, requiring substantially less training data and computational resources than existing methods.

June 21, 2026
StreamKL Delivers up to 43× Speedup in Memory-Efficient Attention Distillation Technology

StreamKL Delivers up to 43× Speedup in Memory-Efficient Attention Distillation

Researchers propose StreamKL, a fused GPU primitive for Kullback-Leibler divergence in attention distillation. It eliminates quadratic memory materialization, enabling up to 43× and 14× speedups in forward and backward passes, and reduces extra HBM footprint to O(1).

June 21, 2026
FastMix: Gradient-Based Data Mixture Optimization Reduces Search Cost in AI Training Technology

FastMix: Gradient-Based Data Mixture Optimization Reduces Search Cost in AI Training

FastMix is a novel framework that automates data mixture discovery by training only a single proxy model and jointly optimizing mixture coefficients and model parameters via gradient descent. It reformulates mixture selection as a bilevel optimization problem, enabling efficient, scalable optimization that outperforms baselines.

June 17, 2026
Learned Image Compression Framework SPARC Boosts VLA Robot Control Performance in Bandwidth-Limited Settings Technology

Learned Image Compression Framework SPARC Boosts VLA Robot Control Performance in Bandwidth-Limited Settings

Researchers introduce SPARC (SPatially Adaptive Rate Control), a learned image compression framework tailored for vision-language-action (VLA) models. SPARC adaptively allocates bitrate based on task relevance and uses a tilted rate loss to preserve critical visual patterns. Experiments on robotic benchmarks RoboCasa365, VLABench, and LIBERO show SPARC achieves stronger control performance than conventional codecs at the same bitrate, with real-world benefits for remote robot control.

June 16, 2026