2026
CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints
AAAI 2026technical
Mixture of Experts (MoE) models have emerged as a promising approach to scale language models efficiently by activating only a subset of parameters for each input. However, deploying these models under GPU memory constraints remains challenging, as existing offloading strategies incur significant ov