2024
Solving Zero-Sum Markov Games with Continuous State via Spectral Dynamic Embedding
NeurIPS 2024poster
In this paper, we propose a provably efficient natural policy gradient algorithm called Spectral Dynamic Embedding Policy Optimization (\SDEPO) for two-player zero-sum stochastic Markov games with continuous state space and finite action space. In the policy evaluation procedure of our algorithm,…