2025
Retraining-free Merging of Sparse MoE via Hierarchical Clustering
ICML 2025poster
Sparse Mixture-of-Experts (SMoE) models represent a significant advancement in large language model (LLM) development through their efficient parameter utilization. These models achieve substantial performance improvements at reduced inference costs. However, the deployment of SMoE models faces cons…