ICLR 2021poster63 citations

Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization

Zhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu, Xiaolong Wang, Fei Fang, Simon Shaolei Du, Yu Wang

Abstract

We propose a simple, general and effective technique, Reward Randomization for discovering diverse strategic policies in complex multi-agent games. Combining reward randomization and policy gradient, we derive a new algorithm, Reward-Randomized Policy Gradient (RPG). RPG is able to discover a set of multiple distinctive human-interpretable strategies in challenging temporal trust dilemmas, including grid-world games and a real-world game Agar.io, where multiple equilibria exist but standard multi-agent policy gradient algorithms always converge to a fixed one with a sub-optimal payoff for every player even using state-of-the-art exploration techniques. Furthermore, with the set of diverse strategies from RPG, we can (1) achieve higher payoffs by fine-tuning the best policy from the set; and (2) obtain an adaptive agent by using this set of strategies as its training opponents.

strategic behaviormulti-agent reinforcement learningreward randomizationdiverse strategies
BibTeX
@inproceedings{
tang2021discovering,
title={Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization},
author={Zhenggang Tang and Chao Yu and Boyuan Chen and Huazhe Xu and Xiaolong Wang and Fei Fang and Simon Shaolei Du and Yu Wang and Yi Wu},
booktitle={International Conference on Learning Representations},
year={2021},
url={https://openreview.net/forum?id=lvRTC669EY_}
}
Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization · ICLR 2021