A Unified Principle of Pessimism for Offline Reinforcement Learning under Model Mismatch
In this paper, we address the challenges of offline reinforcement learning (RL) under model mismatch, where the agent aims to optimize its performance through an offline dataset that may not accurately represent the deployment environment. We identify two primary challenges under the setting: inaccu…