2025
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
NeurIPS 2025poster
Recent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicat…