NeurIPS 2020poster19 citations
Online learning with dynamics: A minimax perspective
Kush Bhatia, Karthik Sridharan
Abstract
We consider the problem of online learning with dynamics, where a learner interacts with a stateful environment over multiple rounds. In each round of the interaction, the learner selects a policy to deploy and incurs a cost that depends on both the chosen policy and current state of the world. The state-evolution dynamics and the costs are allowed to be time-varying, in a possibly adversarial way. In this setting, we study the problem of minimizing policy regret and provide non-constructive upper bounds on the minimax rate for the problem.
BibTeX
@inproceedings{NEURIPS2020_abb451a1,
author = {Bhatia, Kush and Sridharan, Karthik},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
pages = {15020--15030},
publisher = {Curran Associates, Inc.},
title = {Online learning with dynamics: A minimax perspective},
url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/abb451a12cf1a9d93292e81f0d4fdd7a-Paper.pdf},
volume = {33},
year = {2020}
}