2025
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
ICLR 2025poster
Large Language Models (LLMs) with the Mixture-of-Experts (MoE) architectures have shown promising performance on various tasks. However, due to the huge model sizes, running them in resource-constrained environments where the GPU memory is not abundant is challenging. Some existing systems propose t…