← Search

Steven Teig

1 accepted papers

2026

DOT-MoE: Differentiable Optimal Transport for MoEfication

ICML 2026poster

The scaling of Large Language Models (LLMs) has driven significant performance gains but created substantial challenges in inference efficiency. While Mixture of Experts (MoEs) architectures address this by decoupling model size from inference cost, training MoEs from scratch is often unstable and c…

Cited by 0SourceScholar