← Search

Seungki Min

4 accepted papers

2026

Predictive CVaR Q-learning

ICLR 2026poster

We propose a sample-efficient Q-learning algorithm for reinforcement learning with the Conditional Value-at-Risk (CVaR) objective. Our algorithm is built upon predictive tail value function, a novel formulation of risk-sensitive action value, that admits a recursive structure as in the conventional…

Cited by 0SourceScholar
2019

Thompson Sampling with Information Relaxation Penalties

NeurIPS 2019poster

We consider a finite-horizon multi-armed bandit (MAB) problem in a Bayesian setting, for which we propose an information relaxation sampling framework. With this framework, we define an intuitive family of control policies that include Thompson sampling (TS) and the Bayesian optimal policy as endpoi…