← Search

Ben Domingue

2 accepted papers

2026

Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study

AAAI 2026technical

While large language models (LLMs) challenge conventional methods of teaching and learning, they present an exciting opportunity to improve efficiency and scale high-quality instruction. One promising application is the generation of customized exams, tailored to specific course content. There has b

Cited by 0SourcePDFScholar
2026

Noise Tectonics: Measuring the Stability of AI Benchmark Ecosystems

ICML 2026poster

AI benchmark ecosystems compress rich evaluation data into aggregate leaderboard scores, but these scores contain substantial measurement noise whose sources and magnitudes remain unquantified. Without systematic methods to measure this noise and separate signal from artifact, it is unclear when ben…

Cited by 0SourceScholar