2026
DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language Models
AAAI 2026technical
Large Language Models (LLMs) have revolutionized intelligent interactions, enabling mobile applications such as personal assistants on edge devices for local execution. Speculative decoding (SD) has emerged as a promising paradigm to accelerate LLM inference without compromising generation quality,