← Search

Anka Reuel

9 accepted papers

2026

Noise Tectonics: Measuring the Stability of AI Benchmark Ecosystems

ICML 2026poster

AI benchmark ecosystems compress rich evaluation data into aggregate leaderboard scores, but these scores contain substantial measurement noise whose sources and magnitudes remain unquantified. Without systematic methods to measure this noise and separate signal from artifact, it is unclear when ben…

Cited by 0SourceScholar
2026

TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models

ICLR 2026poster

Generative foundation models (GenFMs), such as large language models and text-to-image systems, have demonstrated remarkable capabilities in various downstream applications. As they are increasingly deployed in high-stakes applications, assessing their trustworthiness has become both a critical nece…

Cited by 0SourceScholar
2026

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

ICML 2026poster

Artificial Intelligence (AI) benchmarks play a central role in measuring progress in model development and guiding deployment decisions. However, many benchmarks quickly become saturated, meaning that they can no longer differentiate between the best-performing models, diminishing their long-term va…

Cited by 0SourceScholar
2026

Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations

ICML 2026poster

Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risks and capabilities. Although general capability evaluations are widespread, social impact assessments covering bias, fairness, privacy, environmental costs, a…

Cited by 0SourceScholar
2025

Fantastic Bugs and Where to Find Them in AI Benchmarks

NeurIPS 2025poster

Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a critical bottleneck for reliable evaluation. In this work, we int…

Cited by 0SourceScholar
2025

Position: Human Baselines in Model Evaluations Need Rigor and Transparency (With Recommendations & Reporting Checklist)

ICML 2025spotlight

**In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI performance, and we provide recommendations and a reporting checklist towards this end.** Human performance baselines ar…

2025

Risk Management for Mitigating Benchmark Failure Modes: BenchRisk

NeurIPS 2025poster

Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by various failure modes impacting benchmark bias, variance, coverage, or people's capacity to understand benchmark evidenc…

Cited by 0SourceScholar
2024

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

NeurIPS 2024spotlight

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking progress, and identifying weaknesses in foundation and non-foundati…

Cited by 19SourcePDFScholar
2024

Position: Technical Research and Talent is Needed for Effective AI Governance

ICML 2024oral

In light of recent advancements in AI capabilities and the increasingly widespread integration of AI systems into society, governments worldwide are actively seeking to mitigate the potential harms and risks associated with these technologies through regulation and other governance tools. However, t…

Cited by 6SourcePDFScholar