← Search

Antonie Lin

1 accepted papers

2025

Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention

EMNLP 2025

Speculative decoding is a prominent technique for accelerating LLM inference by leveraging an auxiliary draft model, but its effectiveness is limited by the autoregressive nature of draft generation, where acceptance rates depend on the draft model’s size. Scaling the draft model improves acceptance

Cited by 0SourcePDFScholar