← Search

Stefanos Leonardos

7 accepted papers

2025

Asymptotic Extinction in Large Coordination Games

AAAI 2025technical

We study the exploration-exploitation trade-off for large multiplayer coordination games where players strategise via Q-Learning, a common learning framework in multi-agent reinforcement learning. Q-Learning is known to have two shortcomings, namely non-convergence and potential equilibrium selecti…

Cited by 0SourcePDFScholar
2024

Beating Price of Anarchy and Gradient Descent without Regret in Potential Games

ICLR 2024poster

Arguably one of the thorniest problems in game theory is that of equilibrium selection. Specifically, in the presence of multiple equilibria do self-interested learning dynamics typically select the socially optimal ones? We study a rich class of continuous-time no-regret dynamics in potential games…

Cited by 2SourcePDFScholar
2023

AlberDICE: Addressing Out-Of-Distribution Joint Actions in Offline Multi-Agent RL via Alternating Stationary Distribution Correction Estimation

NeurIPS 2023poster

One of the main challenges in offline Reinforcement Learning (RL) is the distribution shift that arises from the learned policy deviating from the data collection policy. This is often addressed by avoiding out-of-distribution (OOD) actions during policy improvement as their presence can lead to sub…

2022

Global Convergence of Multi-Agent Policy Gradient in Markov Potential Games

ICLR 2022poster

Potential games are arguably one of the most important and widely studied classes of normal form games. They define the archetypal setting of multi-agent coordination in which all agents utilities are perfectly aligned via a common potential function. Can this intuitive framework be transplanted in…

Cited by 160SourcePDFScholar
2021

Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded Rationality

NeurIPS 2021spotlight

The interplay between exploration and exploitation in competitive multi-agent learning is still far from being well understood. Motivated by this, we study smooth Q-learning, a prototypical learning model that explicitly captures the balance between game rewards and exploration costs. We show that Q…

Cited by 39SourcePDFScholar
2021

Exploration-Exploitation in Multi-Agent Learning: Catastrophe Theory Meets Game Theory

AAAI 2021technical

Exploration-exploitation is a powerful and practical tool in multi-agent learning (MAL), however, its effects are far from understood. To make progress in this direction, we study a smooth analogue of Q-learning. We start by showing that our learning model has strong theoretical justification as an…

Cited by 43SourcePDFScholar
2021

Learning in Markets: Greed Leads to Chaos but Following the Price is Right

IJCAI 2021poster

We study learning dynamics in distributed production economies such as blockchain mining, peer-to-peer file sharing and crowdsourcing. These economies can be modelled as multi-product Cournot competitions or all-pay auctions (Tullock contests) when individual firms have market power, or as Fisher ma…

Cited by 26SourcePDFScholar