← Search

Zachary Sunberg

7 accepted papers

2024

Optimality Guarantees for Particle Belief Approximation of POMDPs (Abstract Reprint)

IJCAI 2024poster

Partially observable Markov decision processes (POMDPs) provide a flexible representation for real-world decision and control problems. However, POMDPs are notoriously difficult to solve, especially when the state and observation spaces are continuous or hybrid, which is often the case for physical…

Cited by 0SourcePDFScholar
2024

Recursively-Constrained Partially Observable Markov Decision Processes

UAI 2024poster

Many sequential decision problems involve optimizing one objective function while imposing constraints on other objectives. Constrained Partially Observable Markov Decision Processes (C-POMDP) model this case with transition uncertainty and partial observability. In this work, we first show that C-P…

Cited by 3SourcePDFScholar
2024

Sound Heuristic Search Value Iteration for Undiscounted POMDPs with Reachability Objectives

UAI 2024poster

Partially Observable Markov Decision Processes (POMDPs) are powerful models for sequential decision making under transition and observation uncertainties. This paper studies the challenging yet important problem in POMDPs known as the (indefinite-horizon) Maximal Reachability Probability Problem (MR…

2021

Bayesian Optimized Monte Carlo Planning

AAAI 2021technical

Online solvers for partially observable Markov decision processes have difficulty scaling to problems with large action spaces. Monte Carlo tree search with progressive widening attempts to improve scaling by sampling from the action space to construct a policy search tree. The performance of progre…

2017

Simultaneous active parameter estimation and control using sampling-based Bayesian reinforcement learning

IROS 2017poster

Robots performing manipulation tasks must operate under uncertainty about both their pose and the dynamics of the system. In order to remain robust to modeling error and shifts in payload dynamics, agents must simultaneously perform estimation and control tasks. However, the optimal estimation actio…

Cited by 21SourceScholar