2025
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
NeurIPS 2025poster
Compositional training has been the de-facto paradigm in existing Multimodal Large Language Models (MLLMs), where pre-trained vision encoders are connected with pre-trained LLMs through continuous multimodal pre-training. However, the multimodal scaling property of this paradigm remains difficult…