← Search

Shehreen Azad

4 accepted papers

2025

HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

CVPR 2025poster

Despite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As a result, these models often depend on frame sampling, which risks missing key information over time and lacks task-spec…

Cited by 1SourcePDFScholar