2024
Sequoia: Scalable and Robust Speculative Decoding
NeurIPS 2024spotlight
As the usage of large language models (LLMs) grows, it becomes increasingly important to serve them quickly and efficiently. While speculative decoding has recently emerged as a promising direction for accelerating LLM serving, existing methods are limited in their ability to scale to larger specula…