2026
Reasoning Models Are Test Exploiters: Rethinking Multiple Choice
ICML 2026poster
When evaluating Large Language Models (LLMs) in question-answering domains, multiple-choice question answering (MCQA) is widely used because it enables automatic grading. However, MCQA also exposes models to answer options that can be exploited in ways that inflate reasoning ability. We study this p…