IJCAI 2022poster10 citations

Semi-Supervised Imitation Learning of Team Policies from Suboptimal Demonstrations

Sangwon Seo, Vaibhav V. Unhelkar

Abstract

We present Bayesian Team Imitation Learner (BTIL), an imitation learning algorithm to model the behavior of teams performing sequential tasks in Markovian domains. In contrast to existing multi-agent imitation learning techniques, BTIL explicitly models and infers the time-varying mental states of team members, thereby enabling learning of decentralized team policies from demonstrations of suboptimal teamwork. Further, to allow for sample- and label-efficient policy learning from small datasets, BTIL employs a Bayesian perspective and is capable of learning from semi-supervised demonstrations. We demonstrate and benchmark the performance of BTIL on synthetic multi-agent tasks as well as a novel dataset of human-agent teamwork. Our experiments show that BTIL can successfully learn team policies from demonstrations despite the influence of team members' (time-varying and potentially misaligned) mental states on their behavior.

Humans and AI: Human-AI CollaborationAgent-based and Multi-agent Systems: Human-Agent InteractionMachine Learning: Bayesian LearningRobotics: Human Robot InteractionUncertainty in AI: Sequential Decision Making
BibTeX
@inproceedings{ijcai2022p346,
  title     = {Semi-Supervised Imitation Learning of Team Policies from Suboptimal Demonstrations},
  author    = {Seo, Sangwon and Unhelkar, Vaibhav V.},
  booktitle = {Proceedings of the Thirty-First International Joint Conference on
               Artificial Intelligence, {IJCAI-22}},
  publisher = {International Joint Conferences on Artificial Intelligence Organization},
  editor    = {Lud De Raedt},
  pages     = {2492--2500},
  year      = {2022},
  month     = {7},
  note      = {Main Track},
  doi       = {10.24963/ijcai.2022/346},
  url       = {https://doi.org/10.24963/ijcai.2022/346},
}
Semi-Supervised Imitation Learning of Team Policies from Suboptimal Demonstrations · IJCAI 2022