← Search

Zhengdao Shao

1 accepted papers

2023

Conservative State Value Estimation for Offline Reinforcement Learning

NeurIPS 2023poster

Offline reinforcement learning faces a significant challenge of value over-estimation due to the distributional drift between the dataset and the current learned policy, leading to learning failure in practice. The common approach is to incorporate a penalty term to reward or value estimation in the…