2025
The Emperor’s New Reasoning: Format Imitation Overshadows Genuine Mathematical Understanding in SFT
EMNLP 2025
Recent advances in large language models (LLMs) have yielded impressive gains on mathematical reasoning benchmarks via supervised fine-tuning (SFT). However, the brittleness of these models under input perturbations has cast doubt on whether such improvements reflect genuine reasoning abilities or m