2022
Minimax Optimal Online Imitation Learning via Replay Estimation
NeurIPS 2022accept
Online imitation learning is the problem of how best to mimic expert demonstrations, given access to the environment or an accurate simulator. Prior work has shown that in the \textit{infinite} sample regime, exact moment matching achieves value equivalence to the expert policy. However, in the \tex…