← Search

Yogesh Singh Rawat

7 accepted papers

2025

HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

CVPR 2025poster

Despite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As a result, these models often depend on frame sampling, which risks missing key information over time and lacks task-spec…

Cited by 1SourcePDFScholar
2025

Stable Mean Teacher for Semi-supervised Video Action Detection

AAAI 2025technical

In this work, we focus on semi-supervised learning for video action detection. Video action detection requires spatio-temporal localization in addition to classification, and a limited amount of labels makes the model prone to unreliable predictions. We present Stable Mean Teacher, a simple end-to-e…

2024

Semi-supervised Active Learning for Video Action Detection

AAAI 2024technical

In this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as un- labeled data along with informative sample selection for ac- tion detection. Video action detection requires spatio-te…