← Search

Gong Xudong

2 accepted papers

2024

Iterative Regularized Policy Optimization with Imperfect Demonstrations

ICML 2024poster

Imitation learning heavily relies on the quality of provided demonstrations. In scenarios where demonstrations are imperfect and rare, a prevalent approach for refining policies is through online fine-tuning with reinforcement learning, in which a Kullback–Leibler (KL) regularization is often employ…