← Search

Mulin Li

1 accepted papers

2025

HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference acceleration

EMNLP 2025

Mixture of Experts (MoE) architectures have emerged as a promising paradigm for scaling model capacity through top- k routing mechanisms. Although reducing the number of activated experts inherently enables inference acceleration, this efficiency gain typically comes at the cost of significant perfo