← Search

Aditya Kumar Singh

3 accepted papers

2026

DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference

CVPR 2026

Vision-language models (VLMs) have achieved remarkable multimodal understanding and reasoning capabilities, yet remain computationally expensive due to dense visual tokenization. Existing efficiency approaches either merge redundant visual tokens or drop them progressively in language backbone, ofte

Cited by 0SourcecodeScholar
2023

How You Feelin'? Learning Emotions and Mental States in Movie Scenes

CVPR 2023poster

Movie story analysis requires understanding characters' emotions and mental states. Towards this goal, we formulate emotion understanding as predicting a diverse and multi-label set of emotions at the level of a movie scene and for each character. We propose EmoTx, a multimodal Transformer-based arc…