← Search

Thore Graepel

19 accepted papers

2025

PerturBench: Benchmarking Machine Learning Models for Cellular Perturbation Analysis

NeurIPS 2025poster

We introduce a comprehensive framework for modeling single cell transcriptomic responses to perturbations, aimed at standardizing benchmarking in this rapidly evolving field. Our approach includes a modular and user-friendly model development and evaluation platform, a collection of diverse perturba…

Cited by 0SourcecodeScholar
2022

EigenGame Unloaded: When playing games is better than optimizing

ICLR 2022poster

We build on the recently proposed EigenGame that views eigendecomposition as a competitive game. EigenGame's updates are biased if computed using minibatches of data, which hinders convergence and more sophisticated parallelism in the stochastic setting. In this work, we propose an unbiased stochast…

Cited by 12SourcePDFScholar
2022

NeuPL: Neural Population Learning

ICLR 2022poster

Learning in strategy games (e.g. StarCraft, poker) requires the discovery of diverse policies. This is often achieved by iteratively training new policies against existing ones, growing a policy population that is robust to exploit. This iterative approach suffers from two issues in real-world games…

Cited by 24SourcePDFScholar
2021

A Neural Network Auction For Group Decision Making Over a Continuous Space

IJCAI 2021poster

We propose a system for conducting an auction over locations in a continuous space. It enables participants to express their preferences over possible choices of location in the space, selecting the location that maximizes the total utility of all agents. We prevent agents from tricking the system i…

Cited by 3SourcePDFScholar
2021

Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-Solvers

ICML 2021spotlight

Two-player, constant-sum games are well studied in the literature, but there has been limited progress outside of this setting. We propose Joint Policy-Space Response Oracles (JPSRO), an algorithm for training agents in n-player, general-sum extensive form games, which provably converges to an equil…

2021

Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot

ICML 2021oral

Existing evaluation suites for multi-agent reinforcement learning (MARL) do not assess generalization to novel situations as their primary objective (unlike supervised learning benchmarks). Our contribution, Melting Pot, is a MARL evaluation suite that fills this gap and uses reinforcement learning…

Cited by 117SourcePDFScholar
2020

A Generalized Training Approach for Multiagent Learning

ICLR 2020talk

This paper investigates a population-based training regime based on game-theoretic principles called Policy-Spaced Response Oracles (PSRO). PSRO is general in the sense that it (1) encompasses well-known algorithms such as fictitious play and double oracle as special cases, and (2) in principle appl…

Cited by 127SourcecodeScholar
2020

Learning to Play No-Press Diplomacy with Best Response Policy Iteration

NeurIPS 2020spotlight

Recent advances in deep reinforcement learning (RL) have led to considerable progress in many 2-player zero-sum games, such as Go, Poker and Starcraft. The purely adversarial nature of such games allows for conceptually simple and principled application of RL methods. However real-world settings are…

2020

Smooth markets: A basic mechanism for organizing gradient-based learners

ICLR 2020poster

With the success of modern machine learning, it is becoming increasingly important to understand and control how learning algorithms interact. Unfortunately, negative results from game theory show there is little hope of understanding or controlling general n-player games. We therefore introduce smo…

Cited by 19SourceScholar
2019

Biases for Emergent Communication in Multi-agent Reinforcement Learning

NeurIPS 2019poster

We study the problem of emergent communication, in which language arises because speakers and listeners must communicate information in order to solve tasks. In temporally extended reinforcement learning domains, it has proved hard to learn such communication without centralized training of agents,…

Cited by 99SourcePDFScholar
2019

Emergent Coordination Through Competition

ICLR 2019poster

We study the emergence of cooperative behaviors in reinforcement learning agents by introducing a challenging competitive multi-agent soccer environment with continuous simulated physics. We demonstrate that decentralized, population-based training with co-play can lead to a progression in agents' b…

Cited by 186SourcePDFScholar
2019

Open-ended learning in symmetric zero-sum games

ICML 2019oral

Zero-sum games such as chess and poker are, abstractly, functions that evaluate pairs of agents, for example labeling them ‘winner’ and ‘loser’. If the game is approximately transitive, then self-play generates sequences of agents of increasing strength. However, nontransitive games, such as rock-pa…

Cited by 220SourcePDFScholar
2019

Relational Forward Models for Multi-Agent Learning

ICLR 2019poster

The behavioral dynamics of multi-agent systems have a rich and orderly structure, which can be leveraged to understand these systems, and to improve how artificial agents learn to operate in them. Here we introduce Relational Forward Models (RFM) for multi-agent learning, networks that can learn to…

Cited by 95SourcePDFScholar
2018

Inequity aversion improves cooperation in intertemporal social dilemmas

NeurIPS 2018poster

Groups of humans are often able to find ways to cooperate with one another in complex, temporally extended social dilemmas. Models based on behavioral economics are only able to explain this phenomenon for unrealistic stateless matrix games. Recently, multi-agent reinforcement learning has been appl…

Cited by 296SourcePDFScholar
2018

The Mechanics of n-Player Differentiable Games

ICML 2018oral

The cornerstone underpinning deep learning is the guarantee that gradient descent on an objective converges to local minima. Unfortunately, this guarantee fails in settings, such as generative adversarial nets, where there are multiple interacting losses. The behavior of gradient-based methods in ga…

Cited by 346SourcePDFScholar
2017

A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning

NeurIPS 2017poster

There has been a resurgence of interest in multiagent reinforcement learning (MARL), due partly to the recent success of deep neural networks. The simplest form of MARL is independent reinforcement learning (InRL), where each agent treats all of its experience as part of its (non stationary) environ…

2017

A multi-agent reinforcement learning model of common-pool resource appropriation

NeurIPS 2017poster

Humanity faces numerous problems of common-pool resource appropriation. This class of multi-agent social dilemma includes the problems of ensuring sustainable use of fresh water, common fisheries, grazing pastures, and irrigation systems. Abstract models of common-pool resource appropriation based o…

Cited by 254SourcePDFScholar