2025
Learning Imperfect Information Extensive-form Games with Last-iterate Convergence under Bandit Feedback
ICML 2025poster
We investigate learning approximate Nash equilibrium (NE) policy profiles in two-player zero-sum imperfect information extensive-form games (IIEFGs) with last-iterate convergence guarantees. Existing algorithms either rely on full-information feedback or provide only asymptotic convergence rates. In…