2025
Wait, that’s not an option: LLMs Robustness with Incorrect Multiple-Choice Options
ACL 2025long
This work introduces a novel framework for evaluating LLMs’ capacity to balance instruction-following with critical reasoning when presented with multiple-choice questions containing no valid answers. Through systematic evaluation across arithmetic, domain-specific knowledge, and high-stakes medical…