NeurIPS 2017poster214 citations

QMDP-Net: Deep Learning for Planning under Partial Observability

Peter Karkus, David Hsu, Wee Sun Lee

Abstract

This paper introduces the QMDP-net, a neural network architecture for planning under partial observability. The QMDP-net combines the strengths of model-free learning and model-based planning. It is a recurrent policy network, but it represents a policy for a parameterized set of tasks by connecting a model with a planning algorithm that solves the model, thus embedding the solution structure of planning in a network learning architecture. The QMDP-net is fully differentiable and allows for end-to-end training. We train a QMDP-net on different tasks so that it can generalize to new ones in the parameterized task set and “transfer” to other similar tasks beyond the set. In preliminary experiments, QMDP-net showed strong performance on several robotic tasks in simulation. Interestingly, while QMDP-net encodes the QMDP algorithm, it sometimes outperforms the QMDP algorithm in the experiments, as a result of end-to-end learning.

BibTeX
@inproceedings{NIPS2017_e9412ee5,
 author = {Karkus, Peter and Hsu, David and Lee, Wee Sun},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {I. Guyon and U. Von Luxburg and S. Bengio and H. Wallach and R. Fergus and S. Vishwanathan and R. Garnett},
 pages = {},
 publisher = {Curran Associates, Inc.},
 title = {QMDP-Net: Deep Learning for Planning under Partial Observability},
 url = {https://proceedings.neurips.cc/paper_files/paper/2017/file/e9412ee564384b987d086df32d4ce6b7-Paper.pdf},
 volume = {30},
 year = {2017}
}
QMDP-Net: Deep Learning for Planning under Partial Observability · NeurIPS 2017