2026
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
ICLR 2026oral
Empirical scaling laws have driven the evolution of large language models (LLMs), yet their coefficients shift whenever the model architecture or data pipeline changes. Mixture‑of‑Experts (MoE) models, now standard in state‑of‑the‑art systems, introduce a new sparsity dimension that current dense‑mo…