2023
Toward Human-Like Evaluation for Natural Language Generation with Error Analysis
ACL 2023long
The pretrained language model (PLM) based metrics have been successfully used in evaluating language generation tasks. Recent studies of the human evaluation community show that considering both major errors (e.g. mistranslated tokens) and minor errors (e.g. imperfections in fluency) can produce hig…