2024
A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware Perspective
NeurIPS 2024poster
Offline reinforcement learning endeavors to leverage offline datasets to craft effective agent policy without online interaction, which imposes proper conservative constraints with the support of behavior policies to tackle the out-of-distribution problem. However, existing works often suffer from t…