← Search

Byungjun Lee

2 accepted papers

2021

OptiDICE: Offline Policy Optimization via Stationary Distribution Correction Estimation

ICML 2021spotlight

We consider the offline reinforcement learning (RL) setting where the agent aims to optimize the policy solely from the data without further environment interactions. In offline RL, the distributional shift becomes the primary source of difficulty, which arises from the deviation of the target polic…

Cited by 129SourcePDFScholar
2020

Batch Reinforcement Learning with Hyperparameter Gradients

ICML 2020poster

We consider the batch reinforcement learning problem where the agent needs to learn only from a fixed batch of data, without further interaction with the environment. In such a scenario, we want to prevent the optimized policy from deviating too much from the data collection policy since the estimat…

Cited by 21SourcePDFScholar