ICLR 2018workshop38 citations

Value Propagation Networks

Nantas Nardelli, Gabriel Synnaeve, Zeming Lin, Pushmeet Kohli, Nicolas Usunier

Abstract

We present Value Propagation (VProp), a parameter-efficient differentiable planning module built on Value Iteration which can successfully be trained in a reinforcement learning fashion to solve unseen tasks, has the capability to generalize to larger map sizes, and can learn to navigate in dynamic environments. We evaluate on configurations of MazeBase grid-worlds, with randomly generated environments of several different sizes. Furthermore, we show that the module enables to learn to plan when the environment also includes stochastic elements, providing a cost-efficient learning system to build low-level size-invariant planners for a variety of interactive navigation problems.

Learning to planReinforcement LearningValue IterationNavigationConvnets
BibTeX
@misc{
nardelli2018value,
title={Value Propagation Networks},
author={Nantas Nardelli and Gabriel Synnaeve and Zeming Lin and Pushmeet Kohli and Nicolas Usunier},
year={2018},
url={https://openreview.net/forum?id=Bya8fGWAZ},
}
Value Propagation Networks · ICLR 2018