2026
Flatter Tokens are More Valuable for Speculative Draft Model Training
ICLR 2026poster
Speculative Decoding (SD) is a key technique for accelerating Large Language Model (LLM) inference, but it typically requires training a draft model on a large dataset. We approach this problem from a data-centric perspective, finding that not all training samples contribute equally to the SD accept…