← Search

Mikołaj Zasada

1 accepted papers

2026

SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs

ICML 2026poster

Sparse Mixture-of-Experts (MoE) architectures enable scaling LLM parameters under a fixed inference budget by activating only a small subset of experts via top-k routing. While this preserves causality and suits autoregressive language models, the discrete top-k operator is not differentiable, forci…

Cited by 0SourceScholar