CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints
Mixture of Experts (MoE) models have emerged as a promising approach to scale language models efficiently by activating only a subset of parameters for each input. However, deploying these models under GPU memory constraints remains challenging, as existing offloading strategies incur significant ov