← Search

Daniel Vial

4 accepted papers

2023

Collaborative Multi-Agent Heterogeneous Multi-Armed Bandits

ICML 2023poster

The study of collaborative multi-agent bandits has attracted significant attention recently. In light of this, we initiate the study of a new collaborative setting, consisting of $N$ agents such that each agent is learning one of $M$ stochastic multi-armed bandits to minimize their group cumulative…

Cited by 5SourcePDFScholar
2022

Improved Algorithms for Misspecified Linear Markov Decision Processes

AISTATS 2022poster

For the misspecified linear Markov decision process (MLMDP) model of Jin et al. [2020], we propose an algorithm with three desirable properties. (P1) Its regret after K episodes scales as Kmax{\ensuremath{\varepsilon}mis,\ensuremath{\varepsilon}tol}, where \ensuremath{\varepsilon}mis is the degree o…

Cited by 8SourcePDFScholar
2022

Regret Bounds for Stochastic Shortest Path Problems with Linear Function Approximation

ICML 2022spotlight

We propose an algorithm that uses linear function approximation (LFA) for stochastic shortest path (SSP). Under minimal assumptions, it obtains sublinear regret, is computationally efficient, and uses stationary policies. To our knowledge, this is the first such algorithm in the LFA literature (for…

Cited by 18SourcePDFScholar