2025
A Drop-In Solution for On-the-Fly Adaptation of Speculative Decoding in Large Language Models
ACL 2025long
Large Language Models (LLMs) are cutting-edge generative AI models built on transformer architecture, which tend to be highly memory-intensive when performing real-time inference. Various strategies have been developed to enhance the end-to-end inference speed for LLMs, one of which is speculative d…