← Search

Qingmao Yao

1 accepted papers

2025

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood

ICLR 2025poster

Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the $Q$-value overestimation for out-of-distribution (OOD) actions. Existing methods address this issue by imposing constraints; however, they often become overly conservative when evaluating OOD regions, which cons…