BlonDe: An Automatic Evaluation Metric for Document-level Machine Translation
Yuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell
Abstract
Standard automatic metrics, e.g. BLEU, are not reliable for document-level MT evaluation. They can neither distinguish document-level improvements in translation quality from sentence-level ones, nor identify the discourse phenomena that cause context-agnostic translations. This paper introduces a novel automatic metric BlonDe to widen the scope of automatic MT evaluation from sentence to document level. BlonDe takes discourse coherence into consideration by categorizing discourse-related spans and calculating the similarity-based F1 measure of categorized spans. We conduct extensive comparisons on a newly constructed dataset BWB. The experimental results show that BlonDe possesses better selectivity and interpretability at the document-level, and is more sensitive to document-level nuances. In a large-scale human study, BlonDe also achieves significantly higher Pearson’s r correlation with human judgments compared to previous metrics.
BibTeX
@inproceedings{jiang-etal-2022-blonde,
title = "{BlonDe}: An Automatic Evaluation Metric for Document-level Machine Translation",
author = "Jiang, Yuchen and
Liu, Tianyu and
Ma, Shuming and
Zhang, Dongdong and
Yang, Jian and
Huang, Haoyang and
Sennrich, Rico and
Cotterell, Ryan and
Sachan, Mrinmaya and
Zhou, Ming",
editor = "Carpuat, Marine and
de Marneffe, Marie-Catherine and
Meza Ruiz, Ivan Vladimir",
booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.naacl-main.111/",
doi = "10.18653/v1/2022.naacl-main.111",
pages = "1550--1565"
}