Artificial Intelligence #reinforcement learning#actor-critic
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods
A team of researchers has introduced PAVE (Policy-Aware Value-field Equalization), a critic-centric regularization framework that stabilizes the Q-gradient field in continuous actor-critic reinforcement learning. The method addresses erratic high-frequency oscillations in learned policies without modifying the actor, achieving smoothness comparable to policy-side regularization while maintaining task performance.
Jun 21, 2026 1 source