2024
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
ECCV 2024poster
"Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into Large Language Models (LLMs). The prevailing trend in this field involves the utilization of a vision encoder derived fro…