← Search

Dmitry Sotnikov

2 accepted papers

2024

Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback

ICML 2024poster

In many real-world applications, it is hard to provide a reward signal in each step of a Reinforcement Learning (RL) process and more natural to give feedback when an episode ends. To this end, we study the recently proposed model of RL with Aggregate Bandit Feedback (RL-ABF), where the agent only o…

Cited by 4SourcePDFScholar
2023

Delay-Adapted Policy Optimization and Improved Regret for Adversarial MDP with Delayed Bandit Feedback

ICML 2023poster

Policy Optimization (PO) is one of the most popular methods in Reinforcement Learning (RL). Thus, theoretical guarantees for PO algorithms have become especially important to the RL community. In this paper, we study PO in adversarial MDPs with a challenge that arises in almost every real-world appl…

Cited by 4SourcePDFScholar