← Search

Tal Lancewicki

10 accepted papers

2025

Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback

NeurIPS 2025poster

We study the multi-armed bandit problem with adversarially chosen delays in the Best-of-Both-Worlds (BoBW) framework, which aims to achieve near-optimal performance in both stochastic and adversarial environments. While prior work has made progress toward this goal, existing algorithms suffer from s…

Cited by 0SourceScholar
2025

Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback

ICML 2025poster

We study online finite-horizon Markov Decision Processes with adversarially changing loss and aggregate bandit feedback (a.k.a full-bandit). Under this type of feedback, the agent observes only the total loss incurred over the entire trajectory, rather than the individual losses at each intermediate…

Cited by 1SourcePDFScholar
2023

Delay-Adapted Policy Optimization and Improved Regret for Adversarial MDP with Delayed Bandit Feedback

ICML 2023poster

Policy Optimization (PO) is one of the most popular methods in Reinforcement Learning (RL). Thus, theoretical guarantees for PO algorithms have become especially important to the RL community. In this paper, we study PO in adversarial MDPs with a challenge that arises in almost every real-world appl…

Cited by 4SourcePDFScholar
2023

Regret Minimization and Convergence to Equilibria in General-sum Markov Games

ICML 2023poster

An abundance of recent impossibility results establish that regret minimization in Markov games with adversarial opponents is both statistically and computationally intractable. Nevertheless, none of these results preclude the possibility of regret minimization under the assumption that all parties…

Cited by 30SourcePDFScholar
2022

Learning Adversarial Markov Decision Processes with Delayed Feedback

AAAI 2022technical

Reinforcement learning typically assumes that agents observe feedback for their actions immediately, but in many real-world applications (like recommendation systems) feedback is observed in delay. This paper studies online learning in episodic Markov decision processes (MDPs) with unknown transitio…

Cited by 31SourcePDFScholar
2022

Near-Optimal Regret for Adversarial MDP with Delayed Bandit Feedback

NeurIPS 2022accept

The standard assumption in reinforcement learning (RL) is that agents observe feedback for their actions immediately. However, in practice feedback is often observed in delay. This paper studies online learning in episodic Markov decision process (MDP) with unknown transitions, adversarially changin…

Cited by 26SourcePDFScholar
2021

Stochastic Multi-Armed Bandits with Unrestricted Delay Distributions

ICML 2021spotlight

We study the stochastic Multi-Armed Bandit (MAB) problem with random delays in the feedback received by the algorithm. We consider two settings: the {\it reward dependent} delay setting, where realized delays may depend on the stochastic rewards, and the {\it reward-independent} delay setting. Our m…

Cited by 61SourcePDFScholar