2025
Improved Techniques for Offline Reinforcement Learning: Advantage Value Estimation and Layernorm
ICASSP 2025accepted
Offline reinforcement learning, which aims to learn an optimal policy from a previously collected static datasets. Due to the overestimation caused by extrapolation error, offline algorithms adopt overly pessimistic approaches, which compromise the generalization ability of the learned policy. To ad…