ICLR 2026poster0 citations

Model Predictive Adversarial Imitation Learning for Planning from Observation

Tyler Han, Yanda Bao, Bhaumik Mehta, Gabriel Guo, Sanghun Jung, Anubhav Vishwakarma, Emily Kang, Rosario Scalise

Abstract

Humans can often perform a new task after observing a few demonstrations by inferring the underlying intent. For robots, recovering the intent of the demonstrator through a learned reward function can enable more efficient, interpretable, and robust imitation through planning. A common paradigm for learning how to plan-from-demonstration involves first solving for a reward via Inverse Reinforcement Learning (IRL) and then deploying it via Model Predictive Control (MPC). In this work, we unify these two procedures by introducing planning-based Adversarial Imitation Learning, which simultaneously learns a reward and improves a planning-based agent through experience while using observation-only demonstrations. We study advantages of planning-based AIL in generalization, interpretability, robustness, and sample efficiency through experiments in simulated control tasks and real-world navigation from few or single observation-only demonstration.

Imitation LearningReinforcement LearningModel Predictive Control
BibTeX
@inproceedings{
han2026model,
title={Model Predictive Adversarial Imitation Learning for Planning from Observation},
author={Tyler Han and Yanda Bao and Bhaumik Mehta and Gabriel Guo and Sanghun Jung and Anubhav Vishwakarma and Emily Kang and Rosario Scalise and Jason Liren Zhou and Bryan Xu and Byron Boots},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=rTlPfuKTNg}
}
Model Predictive Adversarial Imitation Learning for Planning from Observation · ICLR 2026