ICASSP 2015accepted0 citations
Risk-averse online learning under mean-variance measures
Abstract
We study risk-averse multi-armed bandit problems under mean-variance measures. We consider two risk mitigation models. In the first model, the variations in the reward values obtained at different times are considered as risk and the objective is to minimize the mean-variance of the observed rewards. In the second model, the quantity of interest is the total reward at the end of the time horizon and the objective is to minimize the mean-variance of the total reward. Under both models, we establish asymptotic as well as finite-time lower bounds on regret and develop online learning a time horizon algorithms that achieve the lower bounds.
BibTeX
@inproceedings{icassp2015_riskaverseonline,
title = {Risk-averse online learning under mean-variance measures},
author = {Sattar Vakili and Qing Zhao},
booktitle = {ICASSP 2015},
year = {2015}
}