2025
Uncertainty-Based Smooth Policy Regularisation for Reinforcement Learning with Few Demonstrations
NeurIPS 2025poster
In reinforcement learning with sparse rewards, demonstrations can accelerate learning, but determining when to imitate them remains challenging. We propose Smooth Policy Regularisation from Demonstrations (SPReD), a framework that addresses the fundamental question: when should an agent imitate a de…