MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional dense models, MoEs achieve better performance with less computation. Speculative decoding (SD) is a widely used techniqu…