← Search

Mehdi Jafarnia-Jahromi

3 accepted papers

2023

Posterior sampling-based online learning for the stochastic shortest path model

UAI 2023poster

We consider the problem of online reinforcement learning for the Stochastic Shortest Path (SSP) problem modeled as an unknown MDP with an absorbing state. We propose PSRL-SSP, a simple posterior sampling-based reinforcement learning algorithm for the SSP problem. The algorithm operates in epochs. At…

Cited by 2SourcePDFScholar
2021

Implicit Finite-Horizon Approximation and Efficient Optimal Algorithms for Stochastic Shortest Path

NeurIPS 2021poster

We introduce a generic template for developing regret minimization algorithms in the Stochastic Shortest Path (SSP) model, which achieves minimax optimal regret as long as certain properties are ensured. The key of our analysis is a new technique called implicit finite-horizon approximation, which a…

Cited by 25SourcePDFScholar
2019

Approximate Relative Value Learning for Average-reward Continuous State MDPs

UAI 2019poster

In this paper, we propose an approximate relative value learning (ARVL) algorithm for non- parametric MDPs with continuous state space and finite actions and average reward criterion. It is a sampling based algorithm combined with kernel density estimation and function approximation via nearest neig…

Cited by 17SourcePDFScholar