ICLR 2026poster0 citations

Negotiated Reasoning: On Provably Addressing Relative Over-Generalization

Junjie Sheng, Yantian Wang, Wenhao Li, Bo Jin, Hongyuan Zha, Jun Wang, Xiangfeng Wang

Abstract

We focus on the relative over-generalization (RO) issue in fully cooperative multi-agent reinforcement learning (MARL). Existing methods show that endowing agents with reasoning can help mitigate RO empirically, but there is little theoretical insight. We first prove that RO is avoided when agents satisfy a consistent reasoning requirement. We then propose a new negotiated reasoning framework connecting reasoning and RO with theoretical guarantees. Based on it, we develop an algorithm called Stein variational negotiated reasoning (SVNR), which uses Stein variational gradient descent to form a negotiation policy that provably bypasses RO under maximum-entropy policy iteration. SVNR is further parameterized with neural networks for computational efficiency. Experiments demonstrate that SVNR significantly outperforms baselines on RO-challenged tasks, including Multi-Agent Particle World and MaMuJoCo, confirming its advantage in achieving better cooperation.

Multi-Agent Reinforcement LearningRelative Over-GeneralizationStein variational gradient descent
BibTeX
@inproceedings{
sheng2026negotiated,
title={Negotiated Reasoning: On Provably Addressing Relative Over-Generalization},
author={Junjie Sheng and Yantian Wang and Wenhao Li and Bo Jin and Hongyuan Zha and Jun Wang and Xiangfeng Wang},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=FmvBrKubtw}
}
Negotiated Reasoning: On Provably Addressing Relative Over-Generalization · ICLR 2026