← Search

Peter Auer

5 accepted papers

2025

Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback

NeurIPS 2025poster

We study the multi-armed bandit problem with adversarially chosen delays in the Best-of-Both-Worlds (BoBW) framework, which aims to achieve near-optimal performance in both stochastic and adversarial environments. While prior work has made progress toward this goal, existing algorithms suffer from s…

Cited by 0SourceScholar
2023

Autonomous Exploration for Navigating in MDPs Using Blackbox RL Algorithms

IJCAI 2023poster

We consider the problem of navigating in a Markov decision process where extrinsic rewards are either absent or ignored. In this setting, the objective is to learn policies to reach all the states that are reachable within a given number of steps (in expectation) from a starting state. We introduce…

Cited by 0SourcePDFScholar
2016

Pareto Front Identification from Stochastic Bandit Feedback

AISTATS 2016poster

We consider the problem of identifying the Pareto front for multiple objectives from a finite set of operating points. Sampling an operating point gives a random vector where each coordinate corresponds to the value of one of the objectives. The Pareto front is the set of operating points that are n…

Cited by 61SourcePDFScholar