← Search

Dane Sherburn

2 accepted papers

2025

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

ICLR 2025oral

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challenging tasks that test real-world ML engineering skills such as training models, pre…

2025

PaperBench: Evaluating AI’s Ability to Replicate AI Research

ICML 2025poster

We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research. Agents must replicate 20 ICML 2024 Spotlight and Oral papers from scratch, including understanding paper contributions, developing a codebase, and successfully executing experiments. F…