2021
Optimizing Percentile Criterion using Robust MDPs
AISTATS 2021poster
We address the problem of computing reliable policies in reinforcement learning problems with limited data. In particular, we compute policies that achieve good returns with high confidence when deployed. This objective, known as the percentile criterion, can be optimized using Robust MDPs (RMDPs).…