← Search

Nils Jansen

23 accepted papers

2026

Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs

IJCAI 2026

We introduce missingness-MDPs (miss-MDPs), a novel subclass of partially observable Markov decision processes (POMDPs) that incorporates the theory of missing data. A miss-MDP is a POMDP whose observation function is a missingness function, specifying the probability that individual state features a

Cited by 0Scholar
2025

Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial Observability

NeurIPS 2025poster

Multi-environment POMDPs (ME-POMDPs) extend standard POMDPs with discrete model uncertainty. ME-POMDPs represent a finite set of POMDPs that share the same state, action, and observation spaces, but may arbitrarily vary in their transition, observation, and reward models. Such models arise, for inst…

Cited by 0SourceScholar
2025

On Evaluating Policies for Robust POMDPs

NeurIPS 2025poster

Robust partially observable Markov decision processes (RPOMDPs) model sequential decision-making problems under partial observability, where an agent must be robust against a range of dynamics. RPOMDPs can be viewed as a two-player game between an agent, who selects actions, and nature, who adversar…

Cited by 0SourceScholar
2025

Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs

IJCAI 2025

Partially observable Markov decision processes (POMDPs) model specific environments in sequential decision-making under uncertainty. Critically, optimal policies for POMDPs may not be robust against perturbations in the environment. Hidden-model POMDPs (HM-POMDPs) capture sets of different environme

Cited by 0SourcePDFScholar
2025

Safety-Prioritizing Curricula for Constrained Reinforcement Learning

ICLR 2025poster

Curriculum learning aims to accelerate reinforcement learning (RL) by generating curricula, i.e., sequences of tasks of increasing difficulty. Although existing curriculum generation approaches provide benefits in sample efficiency, they overlook safety-critical settings where an RL agent must adhe…

Cited by 0SourcePDFScholar
2024

Factored Online Planning in Many-Agent POMDPs

AAAI 2024technical

In centralized multi-agent systems, often modeled as multi-agent partially observable Markov decision processes (MPOMDPs), the action and observation spaces grow exponentially with the number of agents, making the value and belief estimation of single-agent online planning ineffective. Prior work pa…

Cited by 3SourcePDFScholar
2024

Imprecise Probabilities Meet Partial Observability: Game Semantics for Robust POMDPs

IJCAI 2024poster

Partially observable Markov decision processes (POMDPs) rely on the key assumption that probability distributions are precisely known. Robust POMDPs (RPOMDPs) alleviate this concern by defining imprecise probabilities, referred to as uncertainty sets. While robust MDPs have been studied extensively,…

2024

Robust Active Measuring under Model Uncertainty

AAAI 2024technical

Partial observability and uncertainty are common problems in sequential decision-making that particularly impede the use of formal models such as Markov decision processes (MDPs). However, in practice, agents may be able to employ costly sensors to measure their environment and resolve partial obser…

2023

More for Less: Safe Policy Improvement with Stronger Performance Guarantees

IJCAI 2023poster

In an offline reinforcement learning setting, the safe policy improvement (SPI) problem aims to improve the performance of a behavior policy according to which sample data has been generated. State-of-the-art approaches to SPI require a high number of samples to provide practical probabilistic guar…

2023

Probabilities Are Not Enough: Formal Controller Synthesis for Stochastic Dynamical Models with Epistemic Uncertainty

AAAI 2023technical

Capturing uncertainty in models of complex dynamical systems is crucial to designing safe controllers. Stochastic noise causes aleatoric uncertainty, whereas imprecise knowledge of model parameters leads to epistemic uncertainty. Several approaches use formal abstractions to synthesize policies that…

2023

Recursive Small-Step Multi-Agent A* for Dec-POMDPs

IJCAI 2023poster

We present recursive small-step multi-agent A* (RS-MAA*), an exact algorithm that optimizes the expected reward in decentralized partially observable Markov decision processes (Dec-POMDPs). RS-MAA* builds on multi-agent A* (MAA*), an algorithm that finds policies by exploring a search tree, but tack…

Cited by 9SourcePDFScholar
2023

Risk-aware curriculum generation for heavy-tailed task distributions

UAI 2023poster

Automated curriculum generation for reinforcement learning (RL) aims to speed up learning by designing a sequence of tasks of increasing difficulty. Such tasks are usually drawn from probability distributions with exponentially bounded tails, such as uniform or Gaussian distributions. However, exist…

Cited by 3SourcePDFScholar
2023

Safe Reinforcement Learning From Pixels Using a Stochastic Latent Representation

ICLR 2023poster

We address the problem of safe reinforcement learning from pixel observations. Inherent challenges in such settings are (1) a trade-off between reward optimization and adhering to safety constraints, (2) partial observability, and (3) high-dimensional observations. We formalize the problem in a cons…

2023

Safe Reinforcement Learning via Shielding under Partial Observability

AAAI 2023technical

Safe exploration is a common problem in reinforcement learning (RL) that aims to prevent agents from making disastrous decisions while exploring their environment. A family of approaches to this problem assume domain knowledge in the form of a (partial) model of this environment to decide upon the s…

Cited by 54SourcePDFScholar
2022

Robust Anytime Learning of Markov Decision Processes

NeurIPS 2022accept

Markov decision processes (MDPs) are formal models commonly used in sequential decision-making. MDPs capture the stochasticity that may arise, for instance, from imprecise actuators via probabilities in the transition function. However, in data-driven applications, deriving precise probabilities f…

2022

Sampling-Based Robust Control of Autonomous Systems with Non-Gaussian Noise

AAAI 2022technical

Controllers for autonomous systems that operate in safety-critical settings must account for stochastic disturbances. Such disturbances are often modeled as process noise, and common assumptions are that the underlying distributions are known and/or Gaussian. In practice, however, these assumptions…

Cited by 33SourcePDFScholar
2021

Robust Finite-State Controllers for Uncertain POMDPs

AAAI 2021technical

Uncertain partially observable Markov decision processes (uPOMDPs) allow the probabilistic transition and observation functions of standard POMDPs to belong to a so-called uncertainty set. Such uncertainty, referred to as epistemic uncertainty, captures uncountable sets of probability distributions…

Cited by 45SourcePDFScholar
2021

Safe Policies for Factored Partially Observable Stochastic Games

RSS 2021poster

We study planning problems where a controllable agent operates under partial observability and interacts with an uncontrollable opponent; also referred to as the adversary. The agent has two distinct objectives: To maximize an expected value and to adhere to a safety specification. Multi-objective p…

Cited by 8SourcePDFScholar
2020

Robust Policy Synthesis for Uncertain POMDPs via Convex Optimization

IJCAI 2020poster

We study the problem of policy synthesis for uncertain partially observable Markov decision processes (uPOMDPs). The transition probability function of uPOMDPs is only known to belong to a so-called uncertainty set, for instance in the form of probability intervals. Such a model arises when, for e…

Cited by 0SourcePDFScholar