2025
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
EMNLP 2025
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of pr