2025
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
NeurIPS 2025poster
Rapidly improving large language models (LLMs) have the potential to assist in scientific progress. One critical skill in this endeavor is the ability to faithfully reproduce existing work. To evaluate the capability of AI agents to reproduce complex code in an active research area, we introduce the…