2026
Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning
ICML 2026poster
Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement learning (RL) stage. However, SFT is often optimized in isolation to maximize SFT performance alone. We show that, after identical RL training, models initialized…