AAAI 2026technical0 citations

Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement

Jiashu Yao, Heyan Huang, Shuang Zeng, Chuwei Luo, Wangjie You, Jie Tang, Qingsong Liu, Yuhang Guo

Abstract

Through reinforcement learning (RL) with outcome correctness rewards, large reasoning models (LRMs) with scaled inference computation have demonstrated substantial success on complex reasoning tasks. However, the one-sided reward, focused solely on final correctness, limits its ability to provide detailed supervision over internal reasoning process. This deficiency leads to suboptimal internal reasoning quality, manifesting as issues like over-thinking, under-thinking, redundant-thinking, and disordered-thinking. Inspired by the recent progress in LRM self-rewarding, we introduce self-rewriting framework, where a model rewrites its own reasoning texts, and subsequently learns from the rewritten reasoning to improve the internal thought process quality. For algorithm design, we propose a selective rewriting approach wherein only "simple" samples, defined by the model

BibTeX
@inproceedings{aaai2026_incorporatingsel,
  title = {Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement},
  author = {Jiashu Yao and Heyan Huang and Shuang Zeng and Chuwei Luo and Wangjie You and Jie Tang and Qingsong Liu and Yuhang Guo and Yangyang Kang},
  booktitle = {AAAI 2026},
  year = {2026}
}