2026
Sem-MoE: Semantic-aware Model-Data Collaborative Scheduling for Efficient MoE Inference
ICLR 2026poster
Prevailing LLM (Large Language Model) serving engines employ expert parallelism (EP) to implement multi-device inference of massive Mixture-of-Experts (MoE) models. However, the efficiency of expert parallel inference is largely bounded by inter-device communication, as EP embraces expensive all-to-…