Apprenticeship learning in an incompatible feature space
Gakuto Masuyama, Kazunori Umeda
Abstract
This study presents a novel apprenticeship learning method to enable a learner to utilize demonstrations observed in an incompatible feature space. It is assumed that an expert and a learner follow non-identical Markov decision processes (MDPs), and a mapping function is estimated to obtain feature expectation of the demonstrations in an agent space. A conditional density estimation technique is used to represent the feature expectation in closed-form. The proposed method is useful because it is expected to alleviate intractable processes to explicitly specify correspondence of heterogeneous MDPs for apprenticeship learning. Additionally, the method does not require any sampling method to approximate integrals over an agent feature space. A simulation is used to demonstrate the validity of the proposed method in three domains in which it is not possible to directly compare the features of the expert and learner.
BibTeX
@inproceedings{icra2017_apprenticeshiple,
title = {Apprenticeship learning in an incompatible feature space},
author = {Gakuto Masuyama and Kazunori Umeda},
booktitle = {ICRA 2017},
year = {2017}
}