← Search

John Vian

5 accepted papers

2017

Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability

ICML 2017poster

Many real-world tasks involve multiple agents with partial observability and limited communication. Learning is challenging in these settings due to local viewpoints of agents, which perceive the world as non-stationary due to concurrently-exploring teammates. Approaches that learn specialized polic…

Cited by 706SourcePDFScholar
2017

Scalable accelerated decentralized multi-robot policy search in continuous observation spaces

ICRA 2017poster

This paper presents the first ever approach for solving continuous-observation Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs) and their semi-Markovian counterparts, Dec-POSMDPs. This contribution is especially important in robotics, where a vast number of sensors provide c…

Cited by 9SourceScholar
2017

Semantic-level decentralized multi-robot decision-making using probabilistic macro-observations

ICRA 2017poster

Robust environment perception is essential for decision-making on robots operating in complex domains. Intelligent task execution requires principled treatment of uncertainty sources in a robot's observation model. This is important not only for low-level observations (e.g., accelerom-eter data), bu…

Cited by 10SourceScholar
2016

Graph-based Cross Entropy method for solving multi-robot decentralized POMDPs

ICRA 2016

This paper introduces a probabilistic algorithm for multi-robot decision-making under uncertainty, which can be posed as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP). Dec-POMDPs are inherently synchronous decision-making frameworks which require significant computational

Cited by 23SourceScholar
2015

Online heterogeneous multiagent learning under limited communication with applications to forest fire management

IROS 2015poster

Many robotic missions require online estimation of the unknown state transition models associated with uncertainty that stems from mission dynamics. The learning problem is usually distributed among agents in multiagent scenarios, either due to the absence of a centralized processing unit or because…

Cited by 16SourceScholar