AAAI 2026technical0 citations

RLKD: Distilling LLMs’ Reasoning via Reinforcement Learning

Shicheng Xu, Liang Pang, Yunchang Zhu, Jia Gu, Zihao Wei, Jingcheng Deng, Feiyang Pan, Huawei Shen

Abstract

Distilling reasoning paths from teacher to student models via supervised fine-tuning (SFT) provides a shortcut for improving the reasoning ability of the smaller Large Language Models (LLMs). However, the reasoning paths generated by teacher models often reflect only surface-level traces of their underlying authentic reasoning. Insights from cognitive neuroscience suggest that authentic reasoning involves a complex interweaving between meta-reasoning that selects the appropriate sub-problem from multiple candidates, and solving, which addresses the sub-problem. It means that authentic reasoning has implicit multi-branch structure. Supervised fine-tuning collapses this rich structure into a flat sequence of token prediction in teacher

BibTeX
@inproceedings{aaai2026_rlkddistillingll,
  title = {RLKD: Distilling LLMs’ Reasoning via Reinforcement Learning},
  author = {Shicheng Xu and Liang Pang and Yunchang Zhu and Jia Gu and Zihao Wei and Jingcheng Deng and Feiyang Pan and Huawei Shen and Xueqi Cheng},
  booktitle = {AAAI 2026},
  year = {2026}
}
RLKD: Distilling LLMs’ Reasoning via Reinforcement Learning · AAAI 2026