2026
Less Token, More Signal: MoE Expert Pruning via Critical Token Selection
ICML 2026poster
Mixture-of-Experts (MoE) architectures provide strong scalability for large language models, but their large expert parameter footprint poses challenges for efficient deployment. Expert pruning is widely used to reduce model size and inference cost; however, existing approaches are token-agnostic, t…