ICLR 2026poster0 citations

Predictive CVaR Q-learning

Ju-Hyun Kim, Seungki Min

Abstract

We propose a sample-efficient Q-learning algorithm for reinforcement learning with the Conditional Value-at-Risk (CVaR) objective. Our algorithm is built upon predictive tail value function, a novel formulation of risk-sensitive action value, that admits a recursive structure as in the conventional risk-neutral Bellman equation. This structure enables the Q-learning algorithm to utilize the entire set of sample trajectories rather than relying only on worst-case outcomes, enhancing the sample efficiency. We further derive a Bellman optimality equation and a policy improvement theorem, which provide theoretical foundations of our algorithm and remedy inconsistencies that have existed in the literature. Empirical results demonstrate that our method consistently improves CVaR performance while maintaining stable and interpretable learning dynamics.

CVaR optmizationRisk-sensitive RLQ-learningBellman equationPolicy improvemen
BibTeX
@inproceedings{
kim2026predictive,
title={Predictive {CV}aR Q-learning},
author={Ju-Hyun Kim and Seungki Min},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=B4SCegRJOA}
}