← Search

Qingzhi Chen

1 accepted papers

2026

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

ICML 2026poster

Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement learning (RL) stage. However, SFT is often optimized in isolation to maximize SFT performance alone. We show that, after identical RL training, models initialized…

Cited by 0SourceScholar