2020
Metrics also Disagree in the Low Scoring Range: Revisiting Summarization Evaluation Metrics
COLING 2020main
In text summarization, evaluating the efficacy of automatic metrics without human judgments has become recently popular. One exemplar work (Peyrard, 2019) concludes that automatic metrics strongly disagree when ranking high-scoring summaries. In this paper, we revisit their experiments and find that…