2024
MICRO: Model-Based Offline Reinforcement Learning with a Conservative Bellman Operator
IJCAI 2024poster
Offline reinforcement learning (RL) faces a significant challenge of distribution shift. Model-free offline RL penalizes the Q value for out-of-distribution (OOD) data or constrains the policy closed to the behavior policy to tackle this problem, but this inhibits the exploration of the OOD region.…