2024
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
EMNLP 2024main
Large Language Models (LLMs) have become essential in advancing natural language processing (NLP) tasks, but their sequential token generation limits inference speed. Multi-Draft Speculative Decoding (MDSD) offers a promising solution by using a smaller draft model to generate multiple token sequenc…