← Search

Shantanu Jaiswal

3 accepted papers

2024

Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

ICML 2024poster

While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the rich multimodal structures and dynamics from video and text jointly? Or are they achieving high scores by exploiting bia…

Cited by 1SourcePDFScholar
2024

Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios

NeurIPS 2024poster

Complex visual reasoning and question answering (VQA) is a challenging task that requires compositional multi-step processing and higher-level reasoning capabilities beyond the immediate recognition and localization of objects and events. Here, we introduce a fully neural Iterative and Parallel Reas…

Cited by 0SourcePDFScholar
2022

TDAM: Top-Down Attention Module for Contextually Guided Feature Selection in CNNs

ECCV 2022poster

"Attention modules for Convolutional Neural Networks (CNNs) are an effective method to enhance performance on multiple computer-vision tasks. While existing methods appropriately model channel-, spatial- and self-attention, they primarily operate in a feedforward bottom-up manner. Consequently, the…