← Search

A. Bergman

3 accepted papers

2026

Expanding the AI Evaluation Toolbox with Statistical Models

ICML 2026poster

Benchmarks are widely used to evaluate and compare the performance of artificial intelligence systems. However, some approaches to computing benchmark metrics produce invalid uncertainty estimates or make unrecognized assumptions about the evaluation setting. We leverage statistical modeling to make…

Cited by 0SourceScholar
2022

SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems

ACL 2022long

The social impact of natural language processing and its applications has received increasing attention. In this position paper, we focus on the problem of safety for end-to-end conversational AI. We survey the problem landscape therein, introducing a taxonomy of three observed phenomena: the Instig…