← Search

Sara Kangaslahti

4 accepted papers

2026

Boomerang Distillation Enables Zero-Shot Model Size Interpolation

ICLR 2026poster

Large language models (LLMs) are typically deployed under diverse memory and compute constraints. Existing approaches build model families by training each size independently, which is prohibitively expensive and provides only coarse-grained size options. In this work, we identify a novel phenomenon…

Cited by 0SourcecodeScholar