2025
Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention
EMNLP 2025
Speculative decoding is a prominent technique for accelerating LLM inference by leveraging an auxiliary draft model, but its effectiveness is limited by the autoregressive nature of draft generation, where acceptance rates depend on the draft model’s size. Scaling the draft model improves acceptance