2026
Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
ICLR 2026poster
Mixture-of-Experts (MoE) enables efficient scaling of large language models (LLMs) with sparsely activated experts during inference. To effectively deploy large MoE models on memory-constrained devices, many systems introduce expert offloading which caches a subset of experts in fast memory, leaving…