NeurIPS 2025poster0 citations

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements

Bingchen Zhao, Despoina Magka, Minqi Jiang, Xian Li, Roberta Raileanu, Tatiana Shavrina, Jean-Christophe Gagnon-Audet, Kelvin Niu

Abstract

Rapidly improving large language models (LLMs) have the potential to assist in scientific progress. One critical skill in this endeavor is the ability to faithfully reproduce existing work. To evaluate the capability of AI agents to reproduce complex code in an active research area, we introduce the Automated LLM Speedrunning Benchmark, leveraging the research community's contributions to the $\textit{NanoGPT speedrun}$, a competition to train a GPT-2 model in the shortest time. Each of the 19 speedrun tasks provides the agent with the previous record's training script, optionally paired with one of three hint formats, ranging from pseudocode to paper-like descriptions of the new record's improvements. Records execute quickly by design and speedrun improvements encompass diverse code-level changes, ranging from high-level algorithmic advancements to hardware-aware optimizations. These features make the benchmark both accessible and realistic for the frontier problem of improving LLM training. We find that recent frontier reasoning LLMs combined with SoTA scaffolds struggle to reimplement already-known innovations in our benchmark, even when given detailed hints. Our benchmark thus provides a simple, non-saturated measure of an LLM's ability to automate scientific reproduction, a necessary (but not sufficient) skill for an autonomous research agent.

LLM agentsautomated scientific discoveryscientific reproducibilitycode generationAI agent scaffoldingLLM pre-training
BibTeX
@inproceedings{
zhao2025the,
title={The Automated {LLM} Speedrunning Benchmark: Reproducing Nano{GPT} Improvements},
author={Bingchen Zhao and Despoina Magka and Minqi Jiang and Xian Li and Roberta Raileanu and Tatiana Shavrina and Jean-Christophe Gagnon-Audet and Kelvin Niu and Shagun Sodhani and Michael Shvartsman and Andrei Lupu and Alisia Maria Lupidi and Karen Hambardzumyan and Martin Josifoski and Edan Toledo and Thomas Foster and Lucia Cipolina-Kun and Derek Dunfield and Abhishek Charnalia and Alexander H Miller and Oisin Mac Aodha and Jakob Nicolaus Foerster and Yoram Bachrach},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2025},
url={https://openreview.net/forum?id=w98hMEjzu8}
}
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements · NeurIPS 2025