iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing Gold loans jump 93.8% y-o-y, fuel bank credit growth in Q1FY27 Snapchat joins YouTube, LinkedIn and Substack in fight against 'AI slop' Amazon speeds last-mile delivery, expands robotics fleet past 1 million
Home ›› Technology ›› Ai ›› ROSA-RL Uses Reinforcement Learning to Navigate Roundabouts with Uncertainty Awareness

ROSA-RL Uses Reinforcement Learning to Navigate Roundabouts with Uncertainty Awareness

ROSA-RL is an uncertainty-aware speed advisory system for roundabouts that uses reinforcement learning and a Transformer-based model to predict conflict zone occupancy. Evaluated in simulations, it outperforms model-based baselines and nearly matches an ideal scenario with full knowledge.

iG
iGEN Editorial
June 16, 2026
ROSA-RL Uses Reinforcement Learning to Navigate Roundabouts with Uncertainty Awareness

Roundabouts present a major challenge for automated driving because human behavior is heterogeneous and non-deterministic, driving intentions are unknown, and interaction complexity is high. These factors create uncertainty about whether the conflict zone will be blocked or available at the moment of entry. According to a paper published on arXiv, researchers have developed ROSA-RL (Roundabout Optimized Speed Advisory with Reinforcement Learning) to address this problem.

Probabilistic Conflict Forecasting with Transformers

ROSA-RL employs a Transformer-based model to predict conflict zone occupancy over a five-second horizon. The model captures multi-agent interactions, enabling it to anticipate upcoming conflicts and available gaps. According to the paper, the prediction outputs encode uncertainty in future motion and intent, which is then used to augment the state of a classical reinforcement learning (RL) framework. This allows the system to coordinate speed in an uncertainty-aware manner.

Uncertainty-Aware Reinforcement Learning

The core innovation of ROSA-RL is its ability to handle uncertainty explicitly. By incorporating probabilistic conflict forecasts into the RL state representation, the system can make safer and more efficient decisions in mixed traffic environments where human-driven vehicles and automated vehicles interact. The researchers note that this approach closes the gap to an ideal setting that assumes fully known occupancy, while improving both traffic efficiency and safety.

Simulation Evaluation and Results

ROSA-RL was evaluated in simulations grounded in real-world data. According to the paper, the system effectively handles uncertainty and outperforms a comparable model-based baseline. The results demonstrate that the uncertainty-aware RL framework can nearly match the performance of an ideal system with complete knowledge of future occupancy, without requiring that perfect information.

The source code for ROSA-RL is publicly available, as noted in the paper.

Implications for Mixed Traffic Automation

While ROSA-RL is specifically designed for roundabouts, the underlying approach—combining Transformer-based multi-agent prediction with uncertainty-aware reinforcement learning—could be extended to other traffic scenarios involving high interaction complexity and uncertain human behavior. The paper, authored by Schlamp, Anna-Lena; Gerner, Jeremias; Bogenberger, Klaus; Huber, Werner; and Schmidtner, Stefanie, is listed under Computer Science > Artificial Intelligence on arXiv.


Sources:

Keep Reading

Recommended Stories

For the First Time, Zoox Can Charge People for Rides in Its Steering-Wheel-Free Robotaxis Technology

For the First Time, Zoox Can Charge People for Rides in Its Steering-Wheel-Free Robotaxis

The National Highway Traffic Safety Administration (NHTSA) granted Amazon subsidiary Zoox a two-year exemption to deploy up to 5,000 steering-wheel-free robotaxis and charge for rides. Zoox will begin paid operations in Las Vegas, having already transported over 500,000 riders free of charge. The approval marks a milestone for autonomous vehicles built without traditional controls, subject to heightened oversight and safety reporting.

July 30, 2026
Reinforcement Learning Foundation Models: Synthetic MDPs Could Bridge the Gap Technology

Reinforcement Learning Foundation Models: Synthetic MDPs Could Bridge the Gap

The paper by Zighem, Abdelrahman, and Vie argues that reinforcement learning (RL) lacks a foundation model equivalent to those for language and vision. They propose using synthetic Markov Decision Processes (MDPs), which are as feasible to generate as synthetic tabular data, and demonstrate with a Graph Attention Network trained entirely on synthetic MDPs that achieves competitive results without task-specific tuning.

July 8, 2026
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Technology

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation

A new research paper introduces MENTOR, a reinforcement learning framework that uses flexible teacher-optimized rewards to distill tool-use capabilities from large language models into small models. The approach improves out-of-domain generalization compared to supervised fine-tuning and strict reinforcement learning baselines.

July 8, 2026
Free Waymo Rides in California? A Regulatory Quirk Delays Expansion and Keeps Prices at $0 Technology

Free Waymo Rides in California? A Regulatory Quirk Delays Expansion and Keeps Prices at $0

Waymo's new Ojai robotaxis are giving free rides in California because the state's Public Utilities Commission has not yet approved the company's application to expand its service area and add the vehicles to its fleet. The delay, tied to questions about emergency response and unaccompanied minors, means passengers ride without charge while Waymo waits for regulatory sign-off.

July 8, 2026