← Search

Edwin Hamel-De le Court

2 accepted papers

2026

Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning

AAAI 2026technical

Many reinforcement learning algorithms, particularly those that rely on return estimates for policy improvement, can suffer from poor sample efficiency and training instability due to high-variance return estimates. In this paper we leverage new results from off-policy evaluation; it has recently be

Cited by 0SourcePDFScholar
2025

Probabilistic Shielding for Safe Reinforcement Learning

AAAI 2025technical

In real-life scenarios, a Reinforcement Learning (RL) agent aiming to maximize their reward, must often also behave in a safe manner, including at training time. Thus, much attention in recent years has been given to Safe RL, where an agent aims to learn an optimal policy among all policies that sat…

Cited by 0SourcePDFScholar