← Search

Alessio Russo

11 accepted papers

2025

Achieving $\widetilde{\mathcal{O}}(\sqrt{T})$ Regret in Average-Reward POMDPs with Known Observation Models

AISTATS 2025poster

We tackle average-reward infinite-horizon POMDPs with an unknown transition model but a known observation model, a setting that has been previously addressed in two limiting ways: (i) frequentist methods relying on suboptimal stochastic policies having a minimum probability of choosing each action,…

Cited by 0SourceScholar
2023

On the Sample Complexity of Representation Learning in Multi-Task Bandits with Global and Local Structure

AAAI 2023technical

We investigate the sample complexity of learning the optimal arm for multi-task bandit problems. Arms consist of two components: one that is shared across tasks (that we call representation) and one that is task-specific (that we call predictor). The objective is to learn the optimal (representatio…

2021

Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and Learning

NeurIPS 2021spotlight

Importance Sampling (IS) is a widely used building block for a large variety of off-policy estimation and learning algorithms. However, empirical and theoretical studies have progressively shown that vanilla IS leads to poor estimations whenever the behavioral and target policies are too dissimilar.…

2020

Optimal Algorithms for Multiplayer Multi-Armed Bandits

AISTATS 2020poster

The paper addresses various Multiplayer Multi-Armed Bandit (MMAB) problems, where M decision-makers, or players, collaborate to maximize their cumulative reward. We first investigate the MMAB problem where players selecting the same arms experience a collision (and are aware of it) and do not collec…

Cited by 93SourcePDFScholar