2025
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
ICLR 2025poster
Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the $Q$-value overestimation for out-of-distribution (OOD) actions. Existing methods address this issue by imposing constraints; however, they often become overly conservative when evaluating OOD regions, which cons…