NeurIPS 2023poster14 citations

A Finite-Sample Analysis of Payoff-Based Independent Learning in Zero-Sum Stochastic Games

Zaiwei Chen, Kaiqing Zhang, Eric Mazumdar, Asuman E. Ozdaglar, Adam Wierman

Abstract

In this work, we study two-player zero-sum stochastic games and develop a variant of the smoothed best-response learning dynamics that combines independent learning dynamics for matrix games with the minimax value iteration for stochastic games. The resulting learning dynamics are payoff-based, convergent, rational, and symmetric between the two players. Our theoretical results present to the best of our knowledge the first last-iterate finite-sample analysis of such independent learning dynamics. To establish the results, we develop a coupled Lyapunov drift approach to capture the evolution of multiple sets of coupled and stochastic iterates, which might be of independent interest.

Zero-sum stochastic gamespayoff-based independent learningbest-response-type dynamicsfinite-sample analysis
BibTeX
@inproceedings{
chen2023a,
title={A Finite-Sample Analysis of Payoff-Based Independent Learning in Zero-Sum Stochastic Games},
author={Zaiwei Chen and Kaiqing Zhang and Eric Mazumdar and Asuman E. Ozdaglar and Adam Wierman},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems},
year={2023},
url={https://openreview.net/forum?id=hElNdYMs8Z}
}
A Finite-Sample Analysis of Payoff-Based Independent Learning in Zero-Sum Stochastic Games · NeurIPS 2023