MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning
Tristan Tomilin, Luka van den Boogaard, Samuel Garcin, Constantin Ruhdorfer, Bram Grooten, Fabrice Kusters, Yali Du, Andreas Bulling
Abstract
Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. Despite the motivation of lifelong learning, most continual RL papers consider only 3–10 sequential tasks, as CPU-bound environments make longer sequences impractical. Meanwhile, continual learning in cooperative multi-agent settings remains largely unexplored. To address these gaps, we introduce **MEAL** (**M**ulti-agent **E**nvironments for **A**daptive **L**earning), the first benchmark for continual multi-agent RL. By leveraging JAX and GPU acceleration, MEAL enables training on sequences of 100 tasks on a single GPU in a few hours. We find that long task sequences reveal failure modes that do not appear at smaller scales.
BibTeX
@inproceedings{
tomilin2026meal,
title={{MEAL}: A Benchmark for Continual Multi-Agent Reinforcement Learning},
author={Tristan Tomilin and Luka van den Boogaard and Samuel Garcin and Constantin Ruhdorfer and Bram Grooten and Fabrice Kusters and Yali Du and Andreas Bulling and Mykola Pechenizkiy and Meng Fang},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=Mxg6mo1Xzj}
}