2025
Multi-Branch Self-Drafting for LLM Inference Acceleration
AAAI 2025technical
The autoregressive decoding paradigm endows large language models (LLMs) with superior language generation capabilities; however, its step-by-step decoding process inherently limits decoding speed. To mitigate these constraints, the prevalent “draft and validation” strategy enables parallel validati…