← Search

Mehdi Jafarnia Jahromi

4 accepted papers

2024

A Bayesian Learning Algorithm for Unknown Zero-sum Stochastic Games with an Arbitrary Opponent

AISTATS 2024poster

In this paper, we propose Posterior Sampling Reinforcement Learning for Zero-sum Stochastic Games (PSRL-ZSG), the first online learning algorithm that achieves Bayesian regret bound of $\tilde\mathcal{O}(HS\sqrt{AT})$ in the infinite-horizon zero-sum stochastic games with average-reward criterion. H…

Cited by 1SourcePDFScholar
2021

Learning Infinite-horizon Average-reward MDPs with Linear Function Approximation

AISTATS 2021poster

We develop several new algorithms for learning Markov Decision Processes in an infinite-horizon average-reward setting with linear function approximation. Using the optimism principle and assuming that the MDP has a linear structure, we first propose a computationally inefficient algorithm with opti…

Cited by 67SourcePDFScholar
2020

Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes

ICML 2020poster

Model-free reinforcement learning is known to be memory and computation efficient and more amendable to large scale problems. In this paper, two model-free algorithms are introduced for learning infinite-horizon average-reward Markov Decision Processes (MDPs). The first algorithm reduces the problem…

Cited by 135SourcePDFScholar