← Search

Florent Delgrange

4 accepted papers

2024

The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space Models

ICLR 2024poster

Partially Observable Markov Decision Processes (POMDPs) are used to model environments where the state cannot be perceived, necessitating reasoning based on past observations and actions. However, remembering the full history is generally intractable due to the exponential growth in the history spac…

Cited by 7SourcePDFScholar
2023

Wasserstein Auto-encoded MDPs: Formal Verification of Efficiently Distilled RL Policies with Many-sided Guarantees

ICLR 2023poster

Although deep reinforcement learning (DRL) has many success stories, the large-scale deployment of policies learned through these advanced techniques in safety-critical scenarios is hindered by their lack of formal guarantees. Variational Markov Decision Processes (VAE-MDPs) are discrete latent spac…

2022

Distillation of RL Policies with Formal Guarantees via Variational Abstraction of Markov Decision Processes

AAAI 2022technical

We consider the challenge of policy simplification and verification in the context of policies learned through reinforcement learning (RL) in continuous environments. In well-behaved settings, RL algorithms have convergence guarantees in the limit. While these guarantees are valuable, they are insuf…

Cited by 13SourcePDFScholar