← Search

Fabien Pesquerel

3 accepted papers

2023

Fast Asymptotically Optimal Algorithms for Non-Parametric Stochastic Bandits

NeurIPS 2023poster

We consider the problem of regret minimization in non-parametric stochastic bandits. When the rewards are known to be bounded from above, there exists asymptotically optimal algorithms, with asymptotic regret depending on an infimum of Kullback-Leibler divergences (KL). These algorithms are computat…

Cited by 1SourcePDFScholar
2022

IMED-RL: Regret optimal learning of ergodic Markov decision processes

NeurIPS 2022accept

We consider reinforcement learning in a discrete, undiscounted, infinite-horizon Markov decision problem (MDP) under the average reward criterion, and focus on the minimization of the regret with respect to an optimal policy, when the learner does not know the rewards nor transitions of the MDP. In…

Cited by 16SourcePDFScholar
2021

Stochastic bandits with groups of similar arms.

NeurIPS 2021poster

We consider a variant of the stochastic multi-armed bandit problem where arms are known to be organized into different groups having the same mean. The groups are unknown but a lower bound $q$ on their size is known. This situation typically appears when each arm can be described with a list of cate…