NeurIPS 2022accept23 citations
Regret Bounds for Risk-Sensitive Reinforcement Learning
Osbert Bastani, Yecheng Jason Ma, Estelle Shen, Wanqiao Xu
Abstract
In safety-critical applications of reinforcement learning such as healthcare and robotics, it is often desirable to optimize risk-sensitive objectives that account for tail outcomes rather than expected reward. We prove the first regret bounds for reinforcement learning under a general class of risk-sensitive objectives including the popular CVaR objective. Our theory is based on a novel characterization of the CVaR objective as well as a novel optimistic MDP construction.
Risk-sensitive reinforcement learningCVaR objective
BibTeX
@inproceedings{
bastani2022regret,
title={Regret Bounds for Risk-Sensitive Reinforcement Learning},
author={Osbert Bastani and Yecheng Jason Ma and Estelle Shen and Wanqiao Xu},
booktitle={Advances in Neural Information Processing Systems},
editor={Alice H. Oh and Alekh Agarwal and Danielle Belgrave and Kyunghyun Cho},
year={2022},
url={https://openreview.net/forum?id=yJEUDfzsTX7}
}