2025
Beyond Single Frames: Can LMMs Comprehend Implicit Narratives in Comic Strip?
EMNLP 2025
Large Multimodal Models (LMMs) have demonstrated strong performance on vision-language benchmarks, yet current evaluations predominantly focus on single-image reasoning. In contrast, real-world scenarios always involve understanding sequences of images. A typical scenario is comic strips understandi