← Search

Reazul Hasan Russel

2 accepted papers

2021

Optimizing Percentile Criterion using Robust MDPs

AISTATS 2021poster

We address the problem of computing reliable policies in reinforcement learning problems with limited data. In particular, we compute policies that achieve good returns with high confidence when deployed. This objective, known as the percentile criterion, can be optimized using Robust MDPs (RMDPs).…

Cited by 22SourcePDFScholar
2019

Beyond Confidence Regions: Tight Bayesian Ambiguity Sets for Robust MDPs

NeurIPS 2019poster

Robust MDPs (RMDPs) can be used to compute policies with provable worst-case guarantees in reinforcement learning. The quality and robustness of an RMDP solution are determined by the ambiguity set---the set of plausible transition probabilities---which is usually constructed as a multi-dimensional…