← Search

Guillermo Perez

4 accepted papers

2024

The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space Models

ICLR 2024poster

Partially Observable Markov Decision Processes (POMDPs) are used to model environments where the state cannot be perceived, necessitating reasoning based on past observations and actions. However, remembering the full history is generally intractable due to the exponential growth in the history spac…

Cited by 7SourcePDFScholar
2023

Wasserstein Auto-encoded MDPs: Formal Verification of Efficiently Distilled RL Policies with Many-sided Guarantees

ICLR 2023poster

Although deep reinforcement learning (DRL) has many success stories, the large-scale deployment of policies learned through these advanced techniques in safety-critical scenarios is hindered by their lack of formal guarantees. Variational Markov Decision Processes (VAE-MDPs) are discrete latent spac…

2021

Let’s Agree to Degree: Comparing Graph Convolutional Networks in the Message-Passing Framework

ICML 2021oral

In this paper we cast neural networks defined on graphs as message-passing neural networks (MPNNs) to study the distinguishing power of different classes of such models. We are interested in when certain architectures are able to tell vertices apart based on the feature labels given as input with th…

Cited by 43SourcePDFScholar
2020

Finite-Memory Near-Optimal Learning for Markov Decision Processes with Long-Run Average Reward

UAI 2020poster

We consider learning policies online in Markov decision processes with the long-run average reward (a.k.a. mean payoff). To ensure implementability of the policies, we focus on policies with finite memory. Firstly, we show that near optimality can be achieved almost surely, using an unintuitive gadg…

Cited by 8SourcePDFScholar