← Search

Vincent Mai

5 accepted papers

2026

SHAPO: Sharpness-Aware Policy Optimization for Safe Exploration

ICLR 2026poster

Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains. In this paper, we approach safe exploration through the lens of epistemic uncertainty, where the actor’s sensitivity to parameter perturbations serves as a practical proxy for regions of h…

Cited by 0SourceScholar
2025

Safety Representations for Safer Policy Learning

ICLR 2025poster

Reinforcement learning algorithms typically necessitate extensive exploration of the state space to find optimal policies. However, in safety-critical applications, the risks associated with such exploration can lead to catastrophic consequences. Existing safe exploration methods attempt to mitigate…

Cited by 0SourcePDFScholar
2022

Sample Efficient Deep Reinforcement Learning via Uncertainty Estimation

ICLR 2022spotlight

In model-free deep reinforcement learning (RL) algorithms, using noisy value estimates to supervise policy evaluation and optimization is detrimental to the sample efficiency. As this noise is heteroscedastic, its effects can be mitigated using uncertainty-based weights in the optimization process.…

2018

Local Positioning System Using UWB Range Measurements for an Unmanned Blimp

RA-L 2018

Unmanned blimps are a safe and reliable alternative to conventional drones when flying above people. On-board real-time tracking of their pose and velocities is a necessary step toward autonomous navigation. There is a need for an easily deployable technology that is able to accurately and robustly

Cited by 24SourceScholar