2026
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
ICML 2026poster
The massive vocabulary sizes of large language models, often exceeding 100k tokens, impose a computational bottleneck on the final linear projection layer during speculative decoding. Existing vocabulary pruning solutions rely on static or coarsely-grained sub-vocabularies that necessitate large act…