Learning in Zero-Sum Markov Games: Relaxing Strong Reachability and Mixing Time Assumptions
We address payoff-based decentralized learning in infinite-horizon zero-sum Markov games. In this setting, each player makes decisions based solely on received rewards, without observing the opponent