Decentralized TD Tracking with Linear Function Approximation and its Finite-Time Analysis
The present contribution deals with decentralized policy evaluation in multi-agent Markov decision processes using temporal-difference (TD) methods with linear function approximation for scalability. The agents cooperate to estimate the value function of such a process by observing continual state t…