Policy recognition via expectation maximization
Adrian Sosic, Abdelhak M. Zoubir, Heinz Koeppl
Abstract
Learning from Demonstrations (LfD) has proven to be a powerful concept for solving optimal control problems in high-dimensional state spaces where demonstrations can be used to facilitate the search for efficient control policies. However, many existing LfD approaches suffer from either theoretical, practical, or computational drawbacks such as the need to learn a latent reward model, to monitor the expert's controls, or to repeatedly solve potentially demanding planning problems. In this work, we consider the LfD objective from a system identification perspective and propose a probabilistic policy recognition framework based on expectation maximization that operates directly on the observed expert trajectories, avoiding the aforementioned problems. Using a spatial prior over policies, we are able to make accurate predictions in regions of the state space that are scarcely explored.
BibTeX
@inproceedings{icassp2016_policyrecognitio,
title = {Policy recognition via expectation maximization},
author = {Adrian Sosic and Abdelhak M. Zoubir and Heinz Koeppl},
booktitle = {ICASSP 2016},
year = {2016}
}