CasMoE: A Cascaded Framework for Efficient MoE Inference on Resource-constrained Devices
The Mixture-of-Experts (MoE) architecture has emerged as a key enabler for scaling large language models (LLMs), empowering increased model capacity with minimal computational overhead through gating-based dynamic expert activation. However, due to the memory demands introduced by expert modules, Mo