Towards Automated Error Discovery: A Study in Conversational AI
Dominic Petrak, Thy Thy Tran, Iryna Gurevych
Abstract
Although LLM-based conversational agents demonstrate strong fluency and coherence, they still produce undesirable behaviors ( errors ) that are challenging to prevent from reaching users during deployment. Recent research leverages large language models (LLMs) to detect errors and guide response-generation models toward improvement. However, current LLMs struggle to identify errors not explicitly specified in their instructions, such as those arising from updates to the response-generation model or shifts in user behavior. In this work, we introduce Automated Error Discovery , a framework for detecting and defining errors in conversational AI, and propose SEEED (Soft Clustering Extended Encoder-Based Error Detection), as an encoder-based approach to its implementation. We enhance the Soft Nearest Neighbor Loss by amplifying distance weighting for negative samples and introduce Label-Based Sample Ranking to select highly contrastive examples for better representation learning. SEEED outperforms adapted baselines—including GPT-4o and Phi-4—across multiple error-annotated dialogue datasets, improving the accuracy for detecting unknown errors by up to 8 points and demonstrating strong generalization to unknown intent detection.
BibTeX
@inproceedings{emnlp2025_towardsautomated,
title = {Towards Automated Error Discovery: A Study in Conversational AI},
author = {Dominic Petrak and Thy Thy Tran and Iryna Gurevych},
booktitle = {EMNLP 2025},
year = {2025}
}