← Search

Nyle Siddiqui

2 accepted papers

2026

VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues

CVPR 2026

Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual content - actions, objects, and events - directly observable with

Cited by 0SourcecodeScholar
2024

DVANet: Disentangling View and Action Features for Multi-View Action Recognition

AAAI 2024technical

In this work, we present a novel approach to multi-view action recognition where we guide learned action representations to be separated from view-relevant information in a video. When trying to classify action instances captured from multiple viewpoints, there is a higher degree of difficulty due t…