AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
Scalable and reproducible policy evaluation has been a long-standing challenge in robot learning: evaluations are critical to assess progress and build better policies, but evaluation in the real world, especially at a scale that would provide statistically reliable results, is costly in terms of hu…