ICASSP 2021accepted0 citations

On Information Asymmetry in Online Reinforcement Learning

Ezra Tampubolon, Haris Ceribasic, Holger Boche

Abstract

In this work, we study the system of two interacting non-cooperative Q-learning agents, where one agent has the privilege of observing the other's actions. We show that this information asymmetry can lead to a stable outcome of population learning, which does not occur in an environment of general independent learners. Furthermore, we discuss the resulted post-learning policies, show that they are almost optimal in the underlying game sense, and provide numerical hints of almost welfare-optimal of the resulted policies.

BibTeX
@inproceedings{icassp2021_oninformationasy,
  title = {On Information Asymmetry in Online Reinforcement Learning},
  author = {Ezra Tampubolon and Haris Ceribasic and Holger Boche},
  booktitle = {ICASSP 2021},
  year = {2021}
}