← Search

Sebastian Junges

11 accepted papers

2026

Constrained and Robust Policy Synthesis with Satisfiability-Modulo-Probabilistic-Model-Checking

AAAI 2026technical

The ability to compute reward-optimal policies for given and known finite Markov decision processes (MDPs) underpins a variety of applications across planning, controller synthesis, and verification. However, we often want policies (1) to be robust, i.e., they perform well on perturbations of the M

Cited by 0SourcePDFScholar
2025

Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs

IJCAI 2025

Partially observable Markov decision processes (POMDPs) model specific environments in sequential decision-making under uncertainty. Critically, optimal policies for POMDPs may not be robust against perturbations in the environment. Hidden-model POMDPs (HM-POMDPs) capture sets of different environme

Cited by 0SourcePDFScholar
2025

Symbiotic Local Search for Small Decision Tree Policies in MDPs

UAI 2025

We study decision making policies in Markov decision processes (MDPs). Two key performance indicators of such policies are their value and their interpretability. On the one hand, policies that optimize value can be efficiently computed via a plethora of standard methods. However, the representation

Cited by 0SourcePDFScholar
2024

Factored Online Planning in Many-Agent POMDPs

AAAI 2024technical

In centralized multi-agent systems, often modeled as multi-agent partially observable Markov decision processes (MPOMDPs), the action and observation spaces grow exponentially with the number of agents, making the value and belief estimation of single-agent online planning ineffective. Prior work pa…

Cited by 3SourcePDFScholar
2024

Imprecise Probabilities Meet Partial Observability: Game Semantics for Robust POMDPs

IJCAI 2024poster

Partially observable Markov decision processes (POMDPs) rely on the key assumption that probability distributions are precisely known. Robust POMDPs (RPOMDPs) alleviate this concern by defining imprecise probabilities, referred to as uncertainty sets. While robust MDPs have been studied extensively,…

2023

Recursive Small-Step Multi-Agent A* for Dec-POMDPs

IJCAI 2023poster

We present recursive small-step multi-agent A* (RS-MAA*), an exact algorithm that optimizes the expected reward in decentralized partially observable Markov decision processes (Dec-POMDPs). RS-MAA* builds on multi-agent A* (MAA*), an algorithm that finds policies by exploring a search tree, but tack…

Cited by 9SourcePDFScholar
2023

Safe Reinforcement Learning via Shielding under Partial Observability

AAAI 2023technical

Safe exploration is a common problem in reinforcement learning (RL) that aims to prevent agents from making disastrous decisions while exploring their environment. A family of approaches to this problem assume domain knowledge in the form of a (partial) model of this environment to decide upon the s…

Cited by 54SourcePDFScholar
2022

Inductive synthesis of finite-state controllers for POMDPs

UAI 2022poster

We present a novel learning framework to obtain finite-state controllers (FSCs) for partially observable Markov decision processes and illustrate its applicability for indefinite-horizon specifications. Our framework builds on oracle-guided inductive synthesis to explore a design space compactly rep…

2021

Entropy-Guided Control Improvisation

RSS 2021poster

High level declarative constraints provide a powerful (and popular) way to define and construct control policies; however; most synthesis algorithms do not support specifying the degree of randomness (unpredictability) of the resulting controller. In many contexts; e.g.; patrolling; testing; behavio…

Cited by 5SourcePDFScholar
2021

Robust Finite-State Controllers for Uncertain POMDPs

AAAI 2021technical

Uncertain partially observable Markov decision processes (uPOMDPs) allow the probabilistic transition and observation functions of standard POMDPs to belong to a so-called uncertainty set. Such uncertainty, referred to as epistemic uncertainty, captures uncountable sets of probability distributions…

Cited by 45SourcePDFScholar