← Search

Ronald Ortner

8 accepted papers

2023

Autonomous Exploration for Navigating in MDPs Using Blackbox RL Algorithms

IJCAI 2023poster

We consider the problem of navigating in a Markov decision process where extrinsic rewards are either absent or ignored. In this setting, the objective is to learn policies to reach all the states that are reachable within a given number of steps (in expectation) from a starting state. We introduce…

Cited by 0SourcePDFScholar
2019

Regret Bounds for Learning State Representations in Reinforcement Learning

NeurIPS 2019poster

We consider the problem of online reinforcement learning when several state representations (mapping histories to a discrete state space) are available to the learning agent. At least one of these representations is assumed to induce a Markov decision process (MDP), and the performance of the agent…

Cited by 16SourcePDFScholar
2018

Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning

ICML 2018oral

We introduce SCAL, an algorithm designed to perform efficient exploration-exploration in any unknown weakly-communicating Markov Decision Process (MDP) for which an upper bound c on the span of the optimal bias function is known. For an MDP with $S$ states, $A$ actions and $\Gamma \leq S$ possible n…

2016

Improved Learning Complexity in Combinatorial Pure Exploration Bandits

AISTATS 2016poster

We study the problem of combinatorial pure exploration in the stochastic multi-armed bandit problem. We first construct a new measure of complexity that provably characterizes the learning performance of the algorithms we propose for the fixed confidence and the fixed budget setting. We show that th…

Cited by 48SourcePDFScholar
2016

Pareto Front Identification from Stochastic Bandit Feedback

AISTATS 2016poster

We consider the problem of identifying the Pareto front for multiple objectives from a finite set of operating points. Sampling an operating point gives a random vector where each coordinate corresponds to the value of one of the objectives. The Pareto front is the set of operating points that are n…

Cited by 61SourcePDFScholar
2015

Improved Regret Bounds for Undiscounted Continuous Reinforcement Learning

ICML 2015poster

We consider the problem of undiscounted reinforcement learning in continuous state space. Regret bounds in this setting usually hold under various assumptions on the structure of the reward and transition function. Under the assumption that the rewards and transition probabilities are Lipschitz, for…

Cited by 48SourcePDFScholar