← Search

Frederic Thomas Kirstein

2 accepted papers

2025

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator

COLING 2025industry

The quality of meeting summaries generated by natural language generation (NLG) systems is hard to measure automatically. Established metrics such as ROUGE and BERTScore have a relatively low correlation with human judgments and fail to capture nuanced errors. Recent studies suggest using large lang…

2025

What’s Wrong? Refining Meeting Summaries with LLM Feedback

COLING 2025main

Meeting summarization has become a critical task since digital encounters have become a common practice. Large language models (LLMs) show great potential in summarization, offering enhanced coherence and context understanding compared to traditional methods. However, they still struggle to maintain…

Cited by 7SourcePDFScholar