2025
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
ACL 2025finding
Large Language Model (LLM)-based agents have significantly impacted Task-Oriented Dialog Systems (TODS) but continue to face notable performance challenges, especially in zero-shot scenarios. While prior work has noted this performance gap, the behavioral factors driving the performance gap remain u…