ICASSP 2017accepted0 citations

Balancing exploration and exploitation in reinforcement learning using a value of information criterion

Isaac J. Sledge, José C. Príncipe

Abstract

In this paper, we consider an information-theoretic approach for addressing the exploration-exploitation dilemma in reinforcement learning. We employ the value of information, a criterion that provides the optimal trade-off between the expected returns and a policy's degrees of freedom. As the degrees of freedom are reduced, an agent will exploit more than explore. As the policy degrees of freedom increase, an agent will explore more than exploit. We provide an efficient computational procedure for constructing policies using the value of information. The performance is demonstrated on a standard reinforcement learning benchmark problem.

BibTeX
@inproceedings{icassp2017_balancingexplora,
  title = {Balancing exploration and exploitation in reinforcement learning using a value of information criterion},
  author = {Isaac J. Sledge and José C. Príncipe},
  booktitle = {ICASSP 2017},
  year = {2017}
}
Balancing exploration and exploitation in reinforcement learning using a value of information criterion · ICASSP 2017