2025
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
NeurIPS 2025poster
Small language models (SLMs) struggle to learn complex reasoning behaviors, especially when high-quality traces are scarce or difficult to learn from. A typical approach for training such models combines a supervised fine-tuning (SFT) stage, often to distill reasoning capabilities from a larger mode…