2025
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
ICML 2025poster
Behavior regularization, which constrains the policy to stay close to some behavior policy, is widely used in offline reinforcement learning (RL) to manage the risk of hazardous exploitation of unseen actions. Nevertheless, existing literature on behavior-regularized RL primarily focuses on explicit…