2026
Which Reasoning Traces Are Worth Generating Further? Data Curation for Training Reasoning Models
ICML 2026poster
Supervised fine-tuning (SFT) on a small high-quality set of long reasoning traces is an effective way to enable strong reasoning abilities for Large Language Models (LLMs). However, curating a high-quality SFT data requires generating a large pool of long Chain of Thoughts (CoTs), and filtering the …