EMNLP 20250 citations

Language Models Can Easily Learn to Reason from Demonstrations

Dacheng Li, Shiyi Cao, Tyler Griggs, Shu Liu, Xiangxi Mo, Eric Tang, Sumanth Hegde, Kourosh Hakhamaneshi

Abstract

Large reasoning models (LRMs) tackle complex problems by following long chain-of-thoughts (Long CoT) that incorporate reflection, backtracking, and self-validation. However, the training techniques and data requirements to elicit Long CoT remain poorly understood. In this work, we find that language models can effectively learn Long CoT reasoning through data-efficient supervised fine-tuning (SFT) and further parameter-efficient low-rank adaptation (LoRA). Crucially, we find that the structure of Long CoT is critical to the learning process in this data-efficient fine-tuning process. Training on content-incorrect examples, e.g. those lead to incorrect answers or corrupted digits, still leads to significant performance gains. In contrast, training on structurally incorrect examples, e.g., with shuffled or deleted reasoning steps, yield smaller improvements or even degrade performance.

BibTeX
@inproceedings{emnlp2025_languagemodelsca,
  title = {Language Models Can Easily Learn to Reason from Demonstrations},
  author = {Dacheng Li and Shiyi Cao and Tyler Griggs and Shu Liu and Xiangxi Mo and Eric Tang and Sumanth Hegde and Kourosh Hakhamaneshi and Shishir G Patil and Matei Zaharia and Joseph E. Gonzalez and Ion Stoica},
  booktitle = {EMNLP 2025},
  year = {2025}
}
Language Models Can Easily Learn to Reason from Demonstrations · EMNLP 2025