← Search

Hippolyte Bourel

2 accepted papers

2023

Exploration in Reward Machines with Low Regret

AISTATS 2023poster

We study reinforcement learning (RL) for decision processes with non-Markovian reward, in which high-level knowledge in the form of reward machines is available to the learner. Specifically, we investigate the efficiency of RL under the average-reward criterion, in the regret minimization setting. W…

Cited by 11SourcePDFScholar
2020

Tightening Exploration in Upper Confidence Reinforcement Learning

ICML 2020poster

The upper confidence reinforcement learning (UCRL2) algorithm introduced in \citep{jaksch2010near} is a popular method to perform regret minimization in unknown discrete Markov Decision Processes under the average-reward criterion. Despite its nice and generic theoretical regret guarantees, this alg…

Cited by 50SourcePDFScholar