Learning to Play General-Sum Games against Multiple Boundedly Rational Agents
Eric Zhao, Alexander R. Trott, Caiming Xiong, Stephan Zheng
Abstract
We study the problem of training a principal in a multi-agent general-sum game using reinforcement learning (RL). Learning a robust principal policy requires anticipating the worst possible strategic responses of other agents, which is generally NP-hard. However, we show that no-regret dynamics can identify these worst-case responses in poly-time in smooth games. We propose a framework that uses this policy evaluation method for efficiently learning a robust principal policy using RL. This framework can be extended to provide robustness to boundedly rational agents too. Our motivating application is automated mechanism design: we empirically demonstrate our framework learns robust mechanisms in both matrix games and complex spatiotemporal games. In particular, we learn a dynamic tax policy that improves the welfare of a simulated trade-and-barter economy by 15%, even when facing previously unseen boundedly rational RL taxpayers.
BibTeX
@article{Zhao_Trott_Xiong_Zheng_2023, title={Learning to Play General-Sum Games against Multiple Boundedly Rational Agents}, volume={37}, url={https://ojs.aaai.org/index.php/AAAI/article/view/26391}, DOI={10.1609/aaai.v37i10.26391}, abstractNote={We study the problem of training a principal in a multi-agent general-sum game using reinforcement learning (RL). Learning a robust principal policy requires anticipating the worst possible strategic responses of other agents, which is generally NP-hard. However, we show that no-regret dynamics can identify these worst-case responses in poly-time in smooth games. We propose a framework that uses this policy evaluation method for efficiently learning a robust principal policy using RL. This framework can be extended to provide robustness to boundedly rational agents too. Our motivating application is automated mechanism design: we empirically demonstrate our framework learns robust mechanisms in both matrix games and complex spatiotemporal games. In particular, we learn a dynamic tax policy that improves the welfare of a simulated trade-and-barter economy by 15%, even when facing previously unseen boundedly rational RL taxpayers.}, number={10}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Zhao, Eric and Trott, Alexander R. and Xiong, Caiming and Zheng, Stephan}, year={2023}, month={Jun.}, pages={11781-11789} }