2024
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
EMNLP 2024finding
Multimodal Large Language Models (MLLMs) demonstrate a strong understanding of the real world and can even handle complex tasks. However, they still fail on some straightforward visual question-answering (VQA) problems. This paper dives deeper into this issue, revealing that models tend to err when…