2026
Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning
ICLR 2026poster
In this work, we contribute the first approach to solve infinite-horizon discounted general-utility Markov decision processes (GUMDPs) in the single-trial regime, i.e., when the agent's performance is evaluated based on a single trajectory. First, we provide some fundamental results regarding policy…