2021
OptiDICE: Offline Policy Optimization via Stationary Distribution Correction Estimation
ICML 2021spotlight
We consider the offline reinforcement learning (RL) setting where the agent aims to optimize the policy solely from the data without further environment interactions. In offline RL, the distributional shift becomes the primary source of difficulty, which arises from the deviation of the target polic…