2025
HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference acceleration
EMNLP 2025
Mixture of Experts (MoE) architectures have emerged as a promising paradigm for scaling model capacity through top- k routing mechanisms. Although reducing the number of activated experts inherently enables inference acceleration, this efficiency gain typically comes at the cost of significant perfo