2023
Evaluating Open-Domain Question Answering in the Era of Large Language Models
ACL 2023long
Lexical matching remains the de facto evaluation method for open-domain question answering (QA). Unfortunately, lexical matching fails completely when a plausible candidate answer does not appear in the list of gold answers, which is increasingly the case as we shift from extractive to generative mo…