2026
CLIP-FMoE: Scalable CLIP via Fused Mixture-of-Experts with Enforced Specialization
ICLR 2026poster
Mixture-of-Experts (MoE) architectures have emerged as a promising approach for scaling deep learning models while maintaining computational efficiency. However, existing MoE adaptations for Contrastive Language-Image Pre-training (CLIP) models suffer from significant computational overhead during s…