2025
Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked?
ACL 2025finding
Large Language Models (LLMs) have been evaluated using diverse question types, e.g., multiple-choice, true/false, and short/long answers. This study answers an unexplored question about the impact of different question types on LLM accuracy on reasoning tasks. We investigate the performance of five…