← Search

Elizabeth Barnes

3 accepted papers

2025

Measuring AI Ability to Complete Long Software Tasks

NeurIPS 2025poster

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks tha…

Cited by 0SourceScholar
2025

RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts

ICML 2025spotlight

Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations for AI R&D capabilities, and none that are highly realistic and have a direct comparison to human performance. We introduc…

Cited by 16SourcePDFScholar
2023

ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation

NeurIPS 2023oral

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of hi…