← Search

Mohammad Sadegh Talebi

8 accepted papers

2025

Offline RL in Regular Decision Processes: Sample Efficiency via Language Metrics

ICLR 2025poster

This work studies offline Reinforcement Learning (RL) in a class of non-Markovian environments called Regular Decision Processes (RDPs). In RDPs, the unknown dependency of future observations and rewards from the past interactions can be captured by some hidden finite-state automaton. For this reaso…

Cited by 0SourcePDFScholar
2024

Differentially Private No-regret Exploration in Adversarial Markov Decision Processes

UAI 2024poster

We study learning adversarial Markov decision process (MDP) in the episodic setting under the constraint of differential privacy (DP). This is motivated by the widespread applications of reinforcement learning (RL) in non-stationary and even adversarial scenarios, where protecting users’ sensitive i…

Cited by 1SourcePDFScholar
2023

Exploration in Reward Machines with Low Regret

AISTATS 2023poster

We study reinforcement learning (RL) for decision processes with non-Markovian reward, in which high-level knowledge in the form of reward machines is available to the learner. Specifically, we investigate the efficiency of RL under the average-reward criterion, in the regret minimization setting. W…

Cited by 11SourcePDFScholar
2023

Provably Efficient Offline Reinforcement Learning in Regular Decision Processes

NeurIPS 2023poster

This paper deals with offline (or batch) Reinforcement Learning (RL) in episodic Regular Decision Processes (RDPs). RDPs are the subclass of Non-Markov Decision Processes where the dependency on the history of past events can be captured by a finite-state automaton. We consider a setting where the a…

Cited by 5SourcePDFScholar
2020

Adversarial Bandits with Corruptions: Regret Lower Bound and No-regret Algorithm

NeurIPS 2020accepted

This paper studies adversarial bandits with corruptions. In the basic adversarial bandit setting, the reward of arms is predetermined by an adversary who is oblivious to the learner’s policy. In this paper, we consider an extended setting in which an attacker sits in-between the environment and the…

Cited by 42SourcePDFScholar
2020

Tightening Exploration in Upper Confidence Reinforcement Learning

ICML 2020poster

The upper confidence reinforcement learning (UCRL2) algorithm introduced in \citep{jaksch2010near} is a popular method to perform regret minimization in unknown discrete Markov Decision Processes under the average-reward criterion. Despite its nice and generic theoretical regret guarantees, this alg…

Cited by 50SourcePDFScholar