ICASSP 2017accepted0 citations
Balancing exploration and exploitation in reinforcement learning using a value of information criterion
Isaac J. Sledge, José C. Príncipe
Abstract
In this paper, we consider an information-theoretic approach for addressing the exploration-exploitation dilemma in reinforcement learning. We employ the value of information, a criterion that provides the optimal trade-off between the expected returns and a policy's degrees of freedom. As the degrees of freedom are reduced, an agent will exploit more than explore. As the policy degrees of freedom increase, an agent will explore more than exploit. We provide an efficient computational procedure for constructing policies using the value of information. The performance is demonstrated on a standard reinforcement learning benchmark problem.
BibTeX
@inproceedings{icassp2017_balancingexplora,
title = {Balancing exploration and exploitation in reinforcement learning using a value of information criterion},
author = {Isaac J. Sledge and José C. Príncipe},
booktitle = {ICASSP 2017},
year = {2017}
}