RA-L 202415 citations

On the Effectiveness of Retrieval, Alignment, and Replay in Manipulation

Norman Di Palo, Edward Johns

Abstract

Imitation learning with visual observations is notoriously inefficient when addressed with end-to-end behavioural cloning methods. In this letter, we explore an alternative paradigm which decomposes reasoning into three phases. First, a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">retrieval</i> phase, which informs the robot <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">what</i> it can do with an object. Second, an <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">alignment</i> phase, which informs the robot <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">where</i> to interact with the object. And third, a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">replay</i> phase, which informs the robot <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">how</i> to interact with the object. Through a series of real-world experiments on everyday tasks, such as grasping, pouring, and inserting objects, we show that this decomposition brings unprecedented learning efficiency, and effective inter- and intra-class generalisation.

BibTeX
@inproceedings{ral2024_ontheeffectivene,
  title = {On the Effectiveness of Retrieval, Alignment, and Replay in Manipulation},
  author = {Norman Di Palo and Edward Johns},
  booktitle = {RA-L 2024},
  year = {2024}
}