2026
Stabilizing PPO via Latent-Space Regularization and KDE-Driven Exploration
ICML 2026poster
Proximal Policy Optimization (PPO) is widely used in continuous-control tasks, yet its performance is often highly sensitive to training dynamics when neural networks approximate the policy and value functions. This paper introduces SPPO, a drop-in augmentation that preserves PPO’s clipped objective…