← Search

Odalric Maillard

7 accepted papers

2023

Bilinear Exponential Family of MDPs: Frequentist Regret Bound with Tractable Exploration & Planning

AAAI 2023technical

We study the problem of episodic reinforcement learning in continuous state-action spaces with unknown rewards and transitions. Specifically, we consider the setting where the rewards and transitions are modeled using parametric bilinear exponential families. We propose an algorithm, that a) uses pe…

Cited by 15SourcePDFScholar
2021

Optimal Thompson Sampling strategies for support-aware CVaR bandits

ICML 2021spotlight

In this paper we study a multi-arm bandit problem in which the quality of each arm is measured by the Conditional Value at Risk (CVaR) at some level alpha of the reward distribution. While existing works in this setting mainly focus on Upper Confidence Bound algorithms, we introduce a new Thompson S…

2020

Restarted Bayesian Online Change-point Detector achieves Optimal Detection Delay

ICML 2020poster

we consider the problem of sequential change-point detection where both the change-points and the distributions before and after the change are assumed to be unknown. For this problem of primary importance in statistical and sequential learning theory, we derive a variant of the Bayesian Online Cha…

2020

Tightening Exploration in Upper Confidence Reinforcement Learning

ICML 2020poster

The upper confidence reinforcement learning (UCRL2) algorithm introduced in \citep{jaksch2010near} is a popular method to perform regret minimization in unknown discrete Markov Decision Processes under the average-reward criterion. Despite its nice and generic theoretical regret guarantees, this alg…

Cited by 50SourcePDFScholar