2026
Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention
ICML 2026poster
Long-context Large Language Model (LLM) inference has become the norm for today’s AI applications. However, it is severely bottlenecked by the increasing memory demands of its KV cache. Previous works have shown that self-speculative decoding with sparse attention, where tokens are drafted using a s…