← Search

Zeeshan Khan

3 accepted papers

2025

VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

CVPR 2025poster

A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision/video models and a move towards long-video understanding. While exciting, we take a step back and ask: are today's models go…

2024

MICap: A Unified Model for Identity-Aware Movie Descriptions

CVPR 2024poster

Characters are an important aspect of any storyline and identifying and including them in descriptions is necessary for story understanding. While previous work has largely ignored identity and generated captions with someone (anonymized names) recent work formulates id-aware captioning as a fill-in…

Cited by 3SourcePDFScholar