← Search

Lukas Michel

1 accepted papers

2020

Finite-Memory Near-Optimal Learning for Markov Decision Processes with Long-Run Average Reward

UAI 2020poster

We consider learning policies online in Markov decision processes with the long-run average reward (a.k.a. mean payoff). To ensure implementability of the policies, we focus on policies with finite memory. Firstly, we show that near optimality can be achieved almost surely, using an unintuitive gadg…

Cited by 8SourcePDFScholar