ICML 2022oral4 citations
Cooperative Online Learning in Stochastic and Adversarial MDPs
Tal Lancewicki, Aviv Rosenberg, Yishay Mansour
Abstract
We study cooperative online learning in stochastic and adversarial Markov decision process (MDP). That is, in each episode, $m$ agents interact with an MDP simultaneously and share information in order to minimize their individual regret. We consider environments with two types of randomness:
BibTeX
@InProceedings{pmlr-v162-lancewicki22a,
title = {Cooperative Online Learning in Stochastic and Adversarial {MDP}s},
author = {Lancewicki, Tal and Rosenberg, Aviv and Mansour, Yishay},
booktitle = {Proceedings of the 39th International Conference on Machine Learning},
pages = {11918--11968},
year = {2022},
editor = {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
volume = {162},
series = {Proceedings of Machine Learning Research},
month = {17--23 Jul},
publisher = {PMLR},
pdf = {https://proceedings.mlr.press/v162/lancewicki22a/lancewicki22a.pdf},
url = {https://proceedings.mlr.press/v162/lancewicki22a.html},
abstract = {We study cooperative online learning in stochastic and adversarial Markov decision process (MDP). That is, in each episode, $m$ agents interact with an MDP simultaneously and share information in order to minimize their individual regret. We consider environments with two types of randomness: