2021
Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded Rationality
NeurIPS 2021spotlight
The interplay between exploration and exploitation in competitive multi-agent learning is still far from being well understood. Motivated by this, we study smooth Q-learning, a prototypical learning model that explicitly captures the balance between game rewards and exploration costs. We show that Q…