2025
Towards Automated Error Discovery: A Study in Conversational AI
EMNLP 2025
Although LLM-based conversational agents demonstrate strong fluency and coherence, they still produce undesirable behaviors ( errors ) that are challenging to prevent from reaching users during deployment. Recent research leverages large language models (LLMs) to detect errors and guide response-gen