2025
Real-World Summarization: When Evaluation Reaches Its Limits
EMNLP 2025
We examine evaluation of faithfulness to input data in the context of hotel highlights—brief LLM-generated summaries that capture unique features of accommodations. Through human evaluation campaigns involving categorical error assessment and span-level annotation, we compare traditional metrics, tr