Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey
Liangwei Zheng, Wei Emma Zhang, Olaf Maennel, Lin Yue, Weitong Chen
Abstract
Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Despite its growing success, a comprehensive and systematic evaluation of multimodal MoE remains lacking. Existing surveys tend to address either multimodal learning or MoE independently, overlooking the unique interplay between them. This survey fills that gap by addressing a central question: \textit{How does MoE effectively resolve multimodal challenges?} We approach this from three key perspectives: (1) \textbf{MoE as an Efficient Multimodal Framework:} enabling scalable multimodal modeling by decoupling computational cost from parameter growth and mitigating modality redundancy through selective expert activation; (2) \textbf{MoE as a Multimodal Representation Learner:} integrating complementary multi-opinion expert knowledge to enrich alignment and interaction representations; and (3) \textbf{MoE as a Multimodal Adapter:} providing a modular and flexible mechanism to model imperfect modality data such as modality imbalance and missing modality. Through an extensive literature review, we identify critical research gaps, including interpretable routing, expert communication, modality integration, and lifelong multimodal learning. We position this survey as a foundation for future research toward interpretable, adaptive, and sustainable multimodal Mixture-of-Experts systems.
BibTeX
@inproceedings{ijcai2026_tacklingmultimod,
title = {Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey},
author = {Liangwei Zheng and Wei Emma Zhang and Olaf Maennel and Lin Yue and Weitong Chen},
booktitle = {IJCAI 2026},
year = {2026}
}