ICML 2023poster10 citations

Reinforcement Learning with History Dependent Dynamic Contexts

Guy Tennenholtz, Nadav Merlis, Lior Shani, Martin Mladenov, Craig Boutilier

Abstract

We introduce *Dynamic Contextual Markov Decision Processes (DCMDPs)*, a novel reinforcement learning framework for history-dependent environments that generalizes the contextual MDP framework to handle non-Markov environments, where contexts change over time. We consider special cases of the model, with a focus on logistic DCMDPs, which break the exponential dependence on history length by leveraging aggregation functions to determine context transitions. This special structure allows us to derive an upper-confidence-bound style algorithm for which we establish regret bounds. Motivated by our theoretical results, we introduce a practical model-based algorithm for logistic DCMDPs that plans in a latent space and uses optimism over history-dependent features. We demonstrate the efficacy of our approach on a recommendation task (using MovieLens data) where user behavior dynamics evolve in response to recommendations.

BibTeX
@inproceedings{icml2023_reinforcementlea,
  title = {Reinforcement Learning with History Dependent Dynamic Contexts},
  author = {Guy Tennenholtz and Nadav Merlis and Lior Shani and Martin Mladenov and Craig Boutilier},
  booktitle = {ICML 2023},
  year = {2023}
}
Reinforcement Learning with History Dependent Dynamic Contexts · ICML 2023