2026
SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing
AAAI 2026technical
To accelerate Mixture-of-Experts (MoE) inference, the hybrid parallelism paradigm is first applying pipeline parallelism (PP) to vertically divide the model into stages, with each stage further divided horizontally using tensor or expert parallelism. On the algorithm side, dynamic Top-K routing redu