2025
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
NeurIPS 2025spotlight
Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional dense models, MoEs achieve better performance with less computation. Speculative decoding (SD) is a widely used techniqu…