EMNLP 20250 citations

QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering

Woojun Jung, Junyeong Kim

Abstract

Video-to-text summarization remains underexplored in terms of comprehensive evaluation methods. Traditional n-gram overlap-based metrics and recent large language model (LLM)-based approaches depend heavily on human-written reference summaries, limiting their practicality and sensitivity to nuanced semantic aspects. In this paper, we propose QEVA, a reference-free metric evaluating candidate summaries directly against source videos through multimodal question answering. QEVA assesses summaries along three clear dimensions: Coverage, Factuality, and Temporal Coherence. We also introduce MLVU(VS)-Eval, a new annotated benchmark derived from the MLVU dataset, comprising 800 summaries generated from 200 videos using state-of-the-art video-language multimodal models. This dataset establishes a transparent and consistent framework for evaluation. Experimental results demonstrate that QEVA shows higher correlation with human judgments compared to existing approaches, as measured by Kendall’s 𝜏 b , 𝜏 c , and Spearman’s 𝜌 . We hope that our benchmark and metric will facilitate meaningful progress in video-to-text summarization research and provide valuable insights for the development of future evaluation methods.

BibTeX
@inproceedings{emnlp2025_qevaareferencefr,
  title = {QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering},
  author = {Woojun Jung and Junyeong Kim},
  booktitle = {EMNLP 2025},
  year = {2025}
}
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering · EMNLP 2025