← Search

Chloé Rouyer

4 accepted papers

2023

No-Regret Online Reinforcement Learning with Adversarial Losses and Transitions

NeurIPS 2023poster

Existing online learning algorithms for adversarial Markov Decision Processes achieve $\mathcal{O}(\sqrt{T})$ regret after $T$ rounds of interactions even if the loss functions are chosen arbitrarily by an adversary, with the caveat that the transition function has to be fixed. This is because it h…

Cited by 17SourcePDFScholar
2022

A Near-Optimal Best-of-Both-Worlds Algorithm for Online Learning with Feedback Graphs

NeurIPS 2022accept

We consider online learning with feedback graphs, a sequential decision-making framework where the learner's feedback is determined by a directed graph over the action set. We present a computationally-efficient algorithm for learning in this framework that simultaneously achieves near-optimal regre…

Cited by 22SourcePDFScholar
2021

An Algorithm for Stochastic and Adversarial Bandits with Switching Costs

ICML 2021spotlight

We propose an algorithm for stochastic and adversarial multiarmed bandits with switching costs, where the algorithm pays a price $\lambda$ every time it switches the arm being played. Our algorithm is based on adaptation of the Tsallis-INF algorithm of Zimmert and Seldin (2021) and requires no prior…

Cited by 29SourcePDFScholar