2026
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
ICLR 2026poster
Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating research into expert compression. Contrary to recent findings favouring expert *merging* on discriminative benchmarks, we f…