VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To addre…