2023
ACE: Cooperative Multi-Agent Q-learning with Bidirectional Action-Dependency
AAAI 2023technical
Multi-agent reinforcement learning (MARL) suffers from the non-stationarity problem, which is the ever-changing targets at every iteration when multiple agents update their policies at the same time. Starting from first principle, in this paper, we manage to solve the non-stationarity problem by pro…