AAAI 2026technical0 citations

Large Language Models Struggle with Unreasonability in Math Problems

Jingyuan Ma, Damai Dai, Zihang Yuan, Rui Li, Weilin Luo, Bin Wang, Qun Liu, Lei Sha

Abstract

Large Language Models (LLMs) have shown remarkable success on a wide range of math and reasoning benchmarks. However, we observe that they often struggle when faced with unreasonable math problems. Instead of recognizing these issues, models frequently proceed as if the problem is well-posed, producing incorrect answers or falling into overthinking and verbose self-correction. To systematically investigate this overlooked vulnerability, we propose the Unreasonable Math Problems (UMP) benchmark, designed to evaluate LLMs

BibTeX
@inproceedings{aaai2026_largelanguagemod,
  title = {Large Language Models Struggle with Unreasonability in Math Problems},
  author = {Jingyuan Ma and Damai Dai and Zihang Yuan and Rui Li and Weilin Luo and Bin Wang and Qun Liu and Lei Sha and Zhifang Sui},
  booktitle = {AAAI 2026},
  year = {2026}
}
Large Language Models Struggle with Unreasonability in Math Problems · AAAI 2026