← Search

Davide Maran

8 accepted papers

2024

Autoregressive Bandits

AISTATS 2024poster

Autoregressive processes naturally arise in a large variety of real-world scenarios, including stock markets, sales forecasting, weather prediction, advertising, and pricing. When facing a sequential decision-making problem in such a context, the temporal dependence between consecutive observations…

2024

Bandits with Ranking Feedback

NeurIPS 2024poster

In this paper, we introduce a novel variation of multi-armed bandits called bandits with ranking feedback. Unlike traditional bandits, this variation provides feedback to the learner that allows them to rank the arms based on previous pulls, without quantifying numerically the difference in performa…

Cited by 1SourcePDFScholar
2024

Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs

NeurIPS 2024poster

Achieving the no-regret property for Reinforcement Learning (RL) problems in continuous state and action-space environments is one of the major open problems in the field. Existing solutions either work under very specific assumptions or achieve bounds that are vacuous in some regimes. Furthermore,…

Cited by 0SourcePDFScholar
2024

No-Regret Reinforcement Learning in Smooth MDPs

ICML 2024poster

Obtaining no-regret guarantees for reinforcement learning (RL) in the case of problems with continuous state and/or action spaces is still one of the major open challenges in the field. Recently, a variety of solutions have been proposed, but besides very specific settings, the general problem remai…

Cited by 7SourcePDFScholar
2024

Online Markov Decision Processes Configuration with Continuous Decision Space

AAAI 2024technical

In this paper, we investigate the optimal online configuration of episodic Markov decision processes when the space of the possible configurations is continuous. Specifically, we study the interaction between a learner (referred to as the configurator) and an agent with a fixed, unknown policy, when…

Cited by 10SourcePDFScholar
2023

Tight Performance Guarantees of Imitator Policies with Continuous Actions

AAAI 2023technical

Behavioral Cloning (BC) aims at learning a policy that mimics the behavior demonstrated by an expert. The current theoretical understanding of BC is limited to the case of finite actions. In this paper, we study BC with the goal of providing theoretical guarantees on the performance of the imitator…

Cited by 5SourcePDFScholar