2025
Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs
EMNLP 2025
When evaluating large language models (LLMs) with multiple-choice question answering (MCQA), it is common to end the prompt with the string “Answer:” to facilitate automated answer extraction via next-token probabilities. However, there is no consensus on how to tokenize the space following the colo