2025
Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL
ICML 2025poster
Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset. To alleviate extrapolation errors, existing studies often uniformly regularize the value function or policy updates across all states. However, due to substantial variations in data quality, the fixed regula…