DoMoE: Domain-Aware Semantic Expert Prediction for Efficient MoE Inference Under Expert Offloading
Mixture-of-Experts (MoE) large language models improve inference efficiency through sparse expert activation, but deployment on resource-constrained devices remains challenging due to the large expert parameter footprint. Expert offloading mitigates this issue by loading experts on demand, yet its e