ICRA 202553 citations

ICRT: In-Context Imitation Learning via Next-Token Prediction

Max Fu, Huang Huang, Gaurav Datta, Lawrence Yunliang Chen, Will Panitch, Fangchen Liu, Hui Li, Ken Goldberg

Abstract

In-context imitation learning is the capability to perform novel tasks when prompted with task demonstration examples. In-Context Robot Transformer (ICRT) is a causal transformer that performs autoregressive prediction on sensorimotor trajectories, which include images, proprioceptive states, and actions. This approach supports flexible and training-free execution of new tasks at test time. Experiments with a Franka Emika robot demonstrate that ICRT can adapt to new environment configurations that differ from both the prompt and the training data. In a multi-task environment setup, ICRT significantly outperforms current state-of-the-art robot foundation models on generalization to unseen tasks. Code, data, and appendix are available on https://icrt.dev.

BibTeX
@inproceedings{icra2025_icrtincontextimi,
  title = {ICRT: In-Context Imitation Learning via Next-Token Prediction},
  author = {Max Fu and Huang Huang and Gaurav Datta and Lawrence Yunliang Chen and Will Panitch and Fangchen Liu and Hui Li and Ken Goldberg},
  booktitle = {ICRA 2025},
  year = {2025}
}
ICRT: In-Context Imitation Learning via Next-Token Prediction · ICRA 2025