2024
Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exiting
NeurIPS 2024poster
Speculative decoding has demonstrated its effectiveness in accelerating the inference of large language models (LLMs) while maintaining an identical sampling distribution. However, the conventional approach of training separate draft model to achieve a satisfactory token acceptance rate can be costl…