2026
Uncovering the Gradient Geometry of Long CoT: A Spectral-guided Approach to Reasoning Distillation
ICML 2026poster
Large reasoning models (LRMs) achieve remarkable reasoning performance by generating long chains-of-thought (CoT). However, standard supervised fine-tuning (SFT) treats all tokens uniformly, indiscriminately minimizing loss across both essential reasoning steps and those that are noisy, redundant, o…