AAAI 2026technical0 citations

Convergence of Fast Policy Iteration in Markov Games and Robust MDPs

Keith Badger, Jefferson Huang, Marek Petrik

Abstract

Markov games and robust MDPs are closely related models that involve computing a pair of saddle point policies. As part of the long-standing effort to develop efficient algorithms for these models, the Filar-Tolwinski (FT) algorithm has shown considerable promise. As our first contribution, we demonstrate that FT may fail to converge to a saddle point and may loop indefinitely, even in small games. This observation contradicts the proof of FT

BibTeX
@inproceedings{aaai2026_convergenceoffas,
  title = {Convergence of Fast Policy Iteration in Markov Games and Robust MDPs},
  author = {Keith Badger and Jefferson Huang and Marek Petrik},
  booktitle = {AAAI 2026},
  year = {2026}
}
Convergence of Fast Policy Iteration in Markov Games and Robust MDPs · AAAI 2026