2026
Convergence of an actor-critic gradient flow for entropy regularised MDPs in general spaces
ICLR 2026poster
We prove the stability and global convergence of a coupled actor-critic gradient flow for infinite-horizon and entropy-regularised Markov decision processes (MDPs) in continuous state and action space with linear function approximation under Q-function realisability. We consider a version of the act…