← Search

Csaba Szepesvári

7 accepted papers

2024

Exploration via linearly perturbed loss minimisation

AISTATS 2024poster

We introduce \emph{exploration via linear loss perturbations} (EVILL), a randomised exploration method for structured stochastic bandit problems that works by solving for the minimiser of a linearly perturbed regularised negative log-likelihood function. We show that, for the case of generalised lin…

2023

Efficient Planning in Combinatorial Action Spaces with Applications to Cooperative Multi-Agent Reinforcement Learning

AISTATS 2023poster

A practical challenge in reinforcement learning are combinatorial action spaces that make planning computationally demanding. For example, in cooperative multi-agent reinforcement learning, a potentially large number of agents jointly optimize a global reward function, which leads to a combinatorial…

Cited by 5SourcePDFScholar
2022

A free lunch from the noise: Provable and practical exploration for representation learning

UAI 2022poster

Representation learning lies at the heart of the empirical success of deep learning for dealing with the curse of dimensionality. However, the power of representation learning has not been fully exploited yet in reinforcement learning (RL), due to i), the trade-off between expressiveness and tractab…

Cited by 29SourcePDFScholar
2022

Towards painless policy optimization for constrained MDPs

UAI 2022poster

We study policy optimization in an infinite horizon, $\gamma$-discounted constrained Markov decision process (CMDP). Our objective is to return a policy that achieves large expected reward with a small constraint violation. We consider the online setting with linear function approximation and assume…

2019

BubbleRank: Safe Online Learning to Re-Rank via Implicit Click Feedback

UAI 2019poster

In this paper, we study the problem of safe online learning to re-rank, where user feedback is used to improve the quality of displayed lists. Learning to rank has traditionally been studied in two settings. In the offline setting, rankers are typically learned from relevance labels created by judge…

2019

Perturbed-History Exploration in Stochastic Linear Bandits

UAI 2019poster

We propose a new online algorithm for cumulative regret minimization in a stochastic linear bandit. The algorithm pulls the arm with the highest estimated reward in a linear model trained on its perturbed history. Therefore, we call it perturbed-history exploration in a linear bandit (LinPHE). The p…

Cited by 46SourcePDFScholar