Artificial Intelligence #reward#agent
Reward as an Agent: A New Framework for Robust Exploration in Embodied World Models
A new reinforcement learning framework introduces Reward as an Agent to provide robust verification and DynDiff-GRPO for diversified exploration. The method mitigates reward hacking and achieves significant accuracy gains across multiple open-source world models, demonstrating that broader exploration can scale with reliable verification.
Jun 20, 2026 1 source