← Search

Wu Jinzhu

1 accepted papers

2025

Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements

EMNLP 2025

In this paper, we propose a “Generalization Stress Test” to assess Large Language Models’ (LLMs) generalization ability under slight and controlled perturbations, including option length, problem types, and irrelevant noun replacements. We achieve novel and significant findings that, despite high be

Cited by 0SourcePDFScholar