2022
Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation
EMNLP 2022main
The predictions of question answering (QA) systems are typically evaluated against manually annotated finite sets of one or more answers. This leads to a coverage limitation that results in underestimating the true performance of systems, and is typically addressed by extending over exact match (EM)…