2024
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
EMNLP 2024main
Large language models (LLMs) have revolutionized natural language processing and broadened their applicability across diverse commercial applications. However, the deployment of these models is constrained by high inference time in multilingual settings. To mitigate this challenge, this paper explor…