2025
Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference
EMNLP 2025
With the continuous advancement in the performance of large language models (LLMs), their demand for computational resources and memory has significantly increased, which poses major challenges for efficient inference on consumer-grade devices and legacy servers. These devices typically feature rela