2026
Adaptive Scaling of Policy Constraints for Offline Reinforcement Learning
ICLR 2026poster
Offline reinforcement learning (RL) enables learning effective policies from fixed datasets without any environment interaction. Existing methods typically employ policy constraints to mitigate the distribution shift encountered during offline RL training. However, because the scale of the constrain…