← Search

Sean McGregor

6 accepted papers

2025

Lessons for Editors of AI Incidents from the AI Incident Database

AAAI 2025technical

As artificial intelligence (AI) systems become increasingly deployed across the world, they are also increasingly implicated in AI incidents – harm events to individuals and society. As a result, industry, civil society, and governments worldwide are developing best practices and regulations for mon…

2025

Position: In-House Evaluation Is Not Enough. Towards Robust Third-Party Evaluation and Flaw Disclosure for General-Purpose AI

ICML 2025spotlight

The widespread deployment of general-purpose AI (GPAI) systems introduces significant new risks. Yet the infrastructure, practices, and norms for reporting flaws in GPAI systems remain seriously underdeveloped, lagging far behind more established fields like software security. Based on a collaborati…

Cited by 0SourcePDFScholar
2025

Risk Management for Mitigating Benchmark Failure Modes: BenchRisk

NeurIPS 2025poster

Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by various failure modes impacting benchmark bias, variance, coverage, or people's capacity to understand benchmark evidenc…

Cited by 0SourceScholar
2025

To Err Is AI: A Case Study Informing LLM Flaw Reporting Practices

AAAI 2025technical

In August of 2024, 495 hackers generated evaluations in an open-ended bug bounty targeting the Open Language Model (OLMo) from The Allen Institute for AI. A vendor panel staffed by representatives of OLMo's safety program adjudicated changes to OLMo's documentation and awarded cash bounties to parti…

2024

AI Evaluation Authorities: A Case Study Mapping Model Audits to Persistent Standards

AAAI 2024technical

Intelligent system audits are labor-intensive assurance activities that are typically performed once and discarded along with the opportunity to programmatically test all similar products for the market. This study illustrates how several incidents (i.e., harms) involving Named Entity Recognition (N…