2025
Words Over Pixels? Rethinking Vision in Multimodal Large Language Models
IJCAI 2025
Multimodal Large Language Models (MLLMs) promise seamless integration of vision and language understanding. However, despite their strong performance, recent studies reveal that MLLMs often fail to effectively utilize visual information, frequently relying on textual cues instead. This survey provid