← Search

Ronan Fruit

6 accepted papers

2019

Exploration Bonus for Regret Minimization in Discrete and Continuous Average Reward MDPs

NeurIPS 2019poster

The exploration bonus is an effective approach to manage the exploration-exploitation trade-off in Markov Decision Processes (MDPs). While it has been analyzed in infinite-horizon discounted and finite-horizon problems, we focus on designing and analysing the exploration bonus in the more challengin…

2019

Regret Bounds for Learning State Representations in Reinforcement Learning

NeurIPS 2019poster

We consider the problem of online reinforcement learning when several state representations (mapping histories to a discrete state space) are available to the learning agent. At least one of these representations is assumed to induce a Markov decision process (MDP), and the performance of the agent…

Cited by 16SourcePDFScholar
2018

Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning

ICML 2018oral

We introduce SCAL, an algorithm designed to perform efficient exploration-exploration in any unknown weakly-communicating Markov Decision Process (MDP) for which an upper bound c on the span of the optimal bias function is known. For an MDP with $S$ states, $A$ actions and $\Gamma \leq S$ possible n…

2018

Near Optimal Exploration-Exploitation in Non-Communicating Markov Decision Processes

NeurIPS 2018spotlight

While designing the state space of an MDP, it is common to include states that are transient or not reachable by any policy (e.g., in mountain car, the product space of speed and position contains configurations that are not physically reachable). This results in weakly-communicating or multi-chain…

2017

Regret Minimization in MDPs with Options without Prior Knowledge

NeurIPS 2017spotlight

The option framework integrates temporal abstraction into the reinforcement learning model through the introduction of macro-actions (i.e., options). Recent works leveraged on the mapping of Markov decision processes (MDPs) with options to semi-MDPs (SMDPs) and introduced SMDP-versions of exploratio…

Cited by 34SourcePDFScholar