← Search

Heather Frase

2 accepted papers

2025

Lessons for Editors of AI Incidents from the AI Incident Database

AAAI 2025technical

As artificial intelligence (AI) systems become increasingly deployed across the world, they are also increasingly implicated in AI incidents – harm events to individuals and society. As a result, industry, civil society, and governments worldwide are developing best practices and regulations for mon…

2025

Risk Management for Mitigating Benchmark Failure Modes: BenchRisk

NeurIPS 2025poster

Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by various failure modes impacting benchmark bias, variance, coverage, or people's capacity to understand benchmark evidenc…

Cited by 0SourceScholar