← Search

Adrienne Tuynman

2 accepted papers

2024

Finding good policies in average-reward Markov Decision Processes without prior knowledge

NeurIPS 2024poster

We revisit the identification of an $\varepsilon$-optimal policy in average-reward Markov Decision Processes (MDP). In such MDPs, two measures of complexity have appeared in the literature: the diameter, $D$, and the optimal bias span, $H$, which satisfy $H\leq D$. Prior work have studied the comp…

Cited by 3SourcePDFScholar