2024
CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models
ACL 2024long
Multimodal large language models (MLLMs) have demonstrated promising results in a variety of tasks that combine vision and language. As these models become more integral to research and applications, conducting comprehensive evaluations of their capabilities has grown increasingly important. However…