2021
Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors
EMNLP 2021main
Evaluation metrics are a key ingredient for progress of text generation systems. In recent years, several BERT-based evaluation metrics have been proposed (including BERTScore, MoverScore, BLEURT, etc.) which correlate much better with human assessment of text generation quality than BLEU or ROUGE,…