← Search

Inioluwa Deborah Raji

6 accepted papers

2025

From Individual Experience to Collective Evidence: A Reporting-Based Framework for Identifying Systemic Harms

ICML 2025poster

When an individual reports a negative interaction with some system, how can their personal experience be contextualized within broader patterns of system behavior? We study the *reporting database* problem, where individual reports of adverse events arrive sequentially, and are aggregated over time.…

2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

NeurIPS 2025poster

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as `safety' and `robustness' requires strong construct validity, that is, having measures t…

Cited by 0SourceScholar
2025

Position: Medical Large Language Model Benchmarks Should Prioritize Construct Validity

ICML 2025oral

Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician. These claims are usually backed by evaluation on competitive benchmarks—a tradition inherited from mainstream machine learning. But how do we separate real progress…

Cited by 1SourcePDFScholar
2021

AI and the Everything in the Whole Wide World Benchmark

NeurIPS 2021poster

There is a tendency across different subfields in AI to see value in a small collection of influential benchmarks, which we term 'general' benchmarks. These benchmarks operate as stand-ins or abstractions for a range of anointed common problems that are frequently framed as foundational milestones o…

Cited by 344SourceScholar
2021

Are We Learning Yet? A Meta Review of Evaluation Failures Across Machine Learning

NeurIPS 2021poster

Many subfields of machine learning share a common stumbling block: evaluation. Advances in machine learning often evaporate under closer scrutiny or turn out to be less widely applicable than originally hoped. We conduct a meta-review of 107 survey papers from natural language processing, recommen…

Cited by 140SourceScholar