2025
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator
COLING 2025industry
The quality of meeting summaries generated by natural language generation (NLG) systems is hard to measure automatically. Established metrics such as ROUGE and BERTScore have a relatively low correlation with human judgments and fail to capture nuanced errors. Recent studies suggest using large lang…