iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› Llms ›› RollArt: Disaggregated Multi-Task Agentic RL Training at Scale on Alibaba's 3,000-GPU Cluster

RollArt: Disaggregated Multi-Task Agentic RL Training at Scale on Alibaba's 3,000-GPU Cluster

A new system called RollArt, designed for multi-task agentic reinforcement learning training, decouples workload stages onto optimized hardware and achieves 1.31–2.05× training time reduction. It was validated on Alibaba's cluster with over 3,000 GPUs training a hundreds-of-billions-parameter MoE model.

iG
iGEN Editorial
June 17, 2026
RollArt: Disaggregated Multi-Task Agentic RL Training at Scale on Alibaba's 3,000-GPU Cluster

Large-scale training of language models using agentic reinforcement learning (RL) presents a unique challenge: the workload alternates between compute-bound prefill, bandwidth-bound decoding, CPU-heavy environment execution, and bursty reward evaluation. Existing RL systems typically colocate all stages on a single GPU cluster or decouple them only coarsely, leading to hardware underutilization and synchronization overhead. According to a recent paper posted on arXiv, a team of researchers has developed RollArt, a disaggregated system for multi-task agentic RL that maps each pipeline stage to best-fit hardware, dramatically improving throughput and stability.

Disaggregated Architecture

RollArt decouples the RL pipeline at the trajectory level. Instead of synchronizing generation, environment interaction, and reward scoring, the system allows these stages to proceed independently. Specifically:

  • Prefill-heavy tasks are routed to compute-optimized GPUs.
  • Decode-heavy tasks are routed to bandwidth-optimized GPUs.
  • Environment execution is offloaded to CPU clusters.
  • Stateless reward computation is moved to serverless infrastructure.

This disaggregation ensures that slow or failed environment steps never block the other stages. The system also overlaps rollout with training through staleness-bounded asynchronous weight synchronization, eliminating idle time.

Measured Performance Gains

The researchers report that RollArt achieves a 1.31–2.05× reduction in training time compared to various RL systems. The system was evaluated by training a hundreds-of-billions-parameter mixture-of-experts (MoE) model for the Qoder product on an Alibaba cluster with above 3,000 GPUs. The deployment demonstrated both stability and scalability at production scale.

Metric Improvement
Training time reduction 1.31–2.05×
GPU cluster size >3,000 GPUs
Model type MoE (hundreds-of-billions parameters)
Decoupling granularity Trajectory level

Enterprise Implications

For organizations training large-scale AI models, RollArt's disaggregated approach offers a blueprint for reducing infrastructure costs and time-to-deployment. By matching hardware to workload characteristics—compute GPUs for prefill, bandwidth GPUs for decoding, CPU clusters for environments—enterprises can maximize utilization of heterogeneous clusters. The asynchronous training scheme also reduces synchronization penalties, which is critical when training models with hundreds of billions of parameters.

The paper's validation on Alibaba's production environment with thousands of GPUs indicates that the approach is ready for large-scale deployment. While the system is currently demonstrated for agentic RL, the principles of workload disaggregation and asynchronous training could extend to other multi-phase ML workflows.


Sources:

Keep Reading

Recommended Stories

New Robust Q-Learning Algorithm Tackles Mean-Field Control Under Wasserstein Uncertainty Technology

New Robust Q-Learning Algorithm Tackles Mean-Field Control Under Wasserstein Uncertainty

A new robust Q-learning algorithm for discrete-time mean-field control problems under Wasserstein uncertainty in the common noise law combines quantization-and-projection with a Wasserstein dual reformulation. The algorithm, detailed in an arXiv preprint by researchers Laurière, Mathieu, Neufeld, Ariel, Park, and Kyunghyun, establishes convergence with finite-time iteration bounds for both synchronous and asynchronous learning. Numerical experiments on systemic risk and epidemic models illustrate its robustness-performance tradeoff and convergence behavior.

July 8, 2026
New Framework Verifies Safety of Multi-Agent AI Communication for Autonomous Logistics Technology

New Framework Verifies Safety of Multi-Agent AI Communication for Autonomous Logistics

A new framework uses decision tree distillation to formally verify learned communication policies in multi-agent systems, targeting safety-critical autonomous logistics operations. The approach achieves 97.9% fidelity to neural policies and verifies 18 temporal logic properties with 88.9% satisfaction, including collision probabilities below 1% thresholds.

June 22, 2026
Reinforcement-Aware Knowledge Distillation Boosts LLM Reasoning Efficiency Technology

Reinforcement-Aware Knowledge Distillation Boosts LLM Reasoning Efficiency

Researchers propose RL-aware distillation (RLAD) to address distribution mismatch and objective interference in knowledge distillation for LLM reasoning. The method uses Trust Region Ratio Distillation (TRRD) to selectively imitate teacher policies during reinforcement learning. RLAD outperforms offline distillation, standard GRPO, and KL-based on-policy distillation across logic and math benchmarks.

June 21, 2026
CRAX Benchmark Delivers 100x Speedup for Safe Reinforcement Learning Research Technology

CRAX Benchmark Delivers 100x Speedup for Safe Reinforcement Learning Research

Researchers have introduced CRAX (Constrained RL Accelerated with JAX), a fast safe reinforcement learning benchmark that leverages hardware acceleration to achieve up to 100x speedups over CPU-based alternatives. Built on MuJoCo XLA, it includes six environment suites and three agent-specific tasks across three difficulty levels. Evaluation of six popular safe RL methods reveals trade-offs between performance and safety, with curriculum learning improving results.

June 20, 2026