2025
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
ACL 2025finding
This paper introduces the TempVS benchmark, which focuses on temporal grounding and reasoning capabilities of Multimodal Large Language Models (MLLMs) in image sequences. TempVS consists of three main tests (i.e., event relation inference, sentence ordering and image ordering), each accompanied with…