Artificial Intelligence #artificial intelligence#autonomy
Self-Play RL with 30 Minutes of Human Data Trains Coordinated Driving Policies
A new approach from researchers trains autonomous driving policies using self-play reinforcement learning regularized by only 30 minutes of human demonstrations. The method requires 2500x less human data than imitation learning and completes training in 15 hours on a single consumer-grade GPU. The resulting policies successfully coordinate with held-out human trajectories, avoiding the alien driving conventions common in pure self-play systems.
Jun 20, 2026 1 source