M3ARL: Moment-Embedded Mean-Field Multi-Agent Reinforcement Learning for Continuous Action Space
Huaze Tang, Yuanquan Hu, Fanfan Zhao, Junji Yan, Ting Dong, Wenbo Ding
Abstract
Mean-field theory offers a promising solution to the scalability issues encountered in multi-agent reinforcement learning (MARL) within large-scale systems. However, most existing MARL algorithms based on mean-field theory are typically constrained to discrete action space. In continuous action space, the conventional mean-field approximation of mean-field theory breaks down since the representation of mean-field action is not well-defined. In this paper, we propose Moment- Embedded Mean-Field Multi-Agent Reinforcement Learning (M <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> ARL) for continuous action space, embedding mean-field action into multi-order moments. Specifically, we analyze the mean-field approximation on Wasserstein space and derive the approximated form of Q-function. Furthermore, to derive a learn-able form of Q-function, we apply the cylindrical function and propose the moment-embedded representations. Based on the representations and proximal policy optimization (PPO) algorithm, we propose a novel algorithm, M <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> APPO. Finally, we validate the efficacy of the proposed algorithm through experiments of predator-prey environment on continuous action space scenarios.
BibTeX
@inproceedings{icassp2024_m3arlmomentembed,
title = {M3ARL: Moment-Embedded Mean-Field Multi-Agent Reinforcement Learning for Continuous Action Space},
author = {Huaze Tang and Yuanquan Hu and Fanfan Zhao and Junji Yan and Ting Dong and Wenbo Ding},
booktitle = {ICASSP 2024},
year = {2024}
}