← Search

Tristan Tomilin

5 accepted papers

2026

MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning

ICML 2026poster

Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. Despite the motivation of lifelong learning, most continual RL papers consider only 3–10 sequential tasks, as CPU-bound environments make longer sequences impracti…

Cited by 0SourceScholar
2026

Safe Multi-agent Reinforcement Learning with Natural Language Constraints

AAAI 2026technical

Safe Multi-Agent Reinforcement Learning (MARL) typically relies on manually specified numeric cost functions to ensure that policy behaviours respect safety constraints. As systems scale and human-defined constraints become more diverse, context-dependent, and frequently updated, hand-crafting such

Cited by 0SourcePDFScholar
2026

SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas

ICLR 2026poster

Sequential social dilemmas pose a significant challenge in the field of multi-agent reinforcement learning (MARL), requiring environments that accurately reflect the tension between individual and collective interests. Previous benchmarks and environments, such as Melting Pot, provide an evaluation…

Cited by 0SourcecodeScholar
2025

HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents

ICLR 2025poster

Advancing safe autonomous systems through reinforcement learning (RL) requires robust benchmarks to evaluate performance, analyze methods, and assess agent competencies. Humans primarily rely on embodied visual perception to safely navigate and interact with their surroundings, making it a valuable…

Cited by 0SourcePDFScholar
2023

COOM: A Game Benchmark for Continual Reinforcement Learning

NeurIPS 2023poster

The advancement of continual reinforcement learning (RL) has been facing various obstacles, including standardized metrics and evaluation protocols, demanding computational requirements, and a lack of widely accepted standard benchmarks. In response to these challenges, we present COOM ($\textbf{C}$…