2026
Stabilizing Policy Gradient Methods via Reward Profiling
AAAI 2026technical
Policy gradient methods, which have been extensively studied in the last decade, offer an effective and efficient framework for reinforcement learning problems. However, their performances can often be unsatisfactory, suffering from unreliable reward improvements and slow convergence, due to high va