← Search

Ma yinuo

1 accepted papers

2026

The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Reasoning

ICML 2026poster

Knowledge distillation from powerful reasoning models underpins the development of Small Language Models (SLMs). A prevailing assumption in this paradigm is that training data with higher perceived quality, often defined by rigorous logic and superior reward scores, monotonically enhances downstream…

Cited by 0SourceScholar