2025
Exploring the Role of CLIP Global Visual Features in Multimodal Large Language Models
ICASSP 2025accepted
The next recognized development direction of large language models (LLMs) is to integrate and enhance multimodal capability. Although current multimodal large language models (MLLMs) have achieved impressive performance by combining the pre-trained visual encoder CLIP and LLM, these works mainly foc…