ICLR 2026poster0 citations

Sem-MoE: Semantic-aware Model-Data Collaborative Scheduling for Efficient MoE Inference

Zhenyu Zhang, Yan Li, Zhengang Wang, Pengfei chen, Pengfei Zheng

Abstract

Prevailing LLM (Large Language Model) serving engines employ expert parallelism (EP) to implement multi-device inference of massive Mixture-of-Experts (MoE) models. However, the efficiency of expert parallel inference is largely bounded by inter-device communication, as EP embraces expensive all-to-all collectives to route tokens to the remote experts if not collocating on the same GPU/NPU device. Nevertheless, state-of-the-art schemes treat expert device-placement and request (or token) device-scheduling as separate concerns, triggering excessive communication between them and compromising inference efficiency This paper proposes Sem-MoE, a novel \textbf{model-data} collaborative scheduling framework to minimize the steep communication costs in EP-centric MoE serving. Sem-MoE maximally collocates experts and their activating tokens onto the same device using proactively modeled activation likelihood between them and introduces three key techniques: (1) Offline model scheduling, which preliminarily clusters and collocates experts onto devices based on their co-activation tendencies for certain classes of input. (2) Online inter-request data scheduling for Attention-DP setups, which proactively rebatches incoming requests onto the device that hosts experts most likely and frequently activated by the corresponding requests. (3) Online intra-request data scheduling for Attention-TP setups, which seamlessly fuses a token reshuffling procedure into the original inference pipeline and proactively reschedules tokens to devices to reduce dispersed remote routing. We build Sem-MoE into a prevailing LLM serving engine SGLANG. Experiments show our collaborative scheduling approach can effectively reduce the all-to-all communication volume in EP and achieve superior inference throughput compared to existing solutions.

Mixture of ExpertsAll-to-All CommunicationDistributed Inference
BibTeX
@inproceedings{
zhang2026semmoe,
title={Sem-MoE: Semantic-aware Model-Data Collaborative Scheduling for Efficient MoE Inference},
author={Zhenyu Zhang and Yan Li and Zhengang Wang and Pengfei chen and Pengfei Zheng},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=MSHPrMpIHZ}
}