UAI 2021poster4 citations
No-regret learning with high-probability in adversarial Markov decision processes
Mahsa Ghasemi, Abolfazl Hashemi, Haris Vikalo, Ufuk Topcu
Abstract
In a variety of problems, a decision-maker is unaware of the loss function associated with a task, yet it has to minimize this unknown loss in order to accomplish the task. Furthermore, the decision-maker’s task may evolve, resulting in a varying loss function. In this setting, we explore sequential decision-making problems modeled by adversarial Markov decision processes, where the loss function may arbitrarily change at every time step. We consider the bandit feedback scenario, where the agent observes only the loss corresponding to its actions. We propose an algorithm, called
BibTeX
@InProceedings{pmlr-v161-ghasemi21a,
title = {No-regret learning with high-probability in adversarial Markov decision processes},
author = {Ghasemi, Mahsa and Hashemi, Abolfazl and Vikalo, Haris and Topcu, Ufuk},
booktitle = {Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence},
pages = {992--1001},
year = {2021},
editor = {de Campos, Cassio and Maathuis, Marloes H.},
volume = {161},
series = {Proceedings of Machine Learning Research},
month = {27--30 Jul},
publisher = {PMLR},
pdf = {https://proceedings.mlr.press/v161/ghasemi21a/ghasemi21a.pdf},
url = {https://proceedings.mlr.press/v161/ghasemi21a.html},
abstract = {In a variety of problems, a decision-maker is unaware of the loss function associated with a task, yet it has to minimize this unknown loss in order to accomplish the task. Furthermore, the decision-maker’s task may evolve, resulting in a varying loss function. In this setting, we explore sequential decision-making problems modeled by adversarial Markov decision processes, where the loss function may arbitrarily change at every time step. We consider the bandit feedback scenario, where the agent observes only the loss corresponding to its actions. We propose an algorithm, called