How Generalizable is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation
With the rise of stochastic generative models in robot policy learning, end-to-end visuomotor policies are increasingly successful at solving complex tasks by learning from human demonstrations. Nevertheless, since real-world evaluation costs afford users only a small number of policy rollouts, it r