← Search

Rohit Gupta

10 accepted papers

2026

VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues

CVPR 2026

Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual content - actions, objects, and events - directly observable with

Cited by 0SourcecodeScholar
2026

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale

CVPR 2026

The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensics, social media, and exploration. Existing classification-based approaches operate at a coarse city-level granularity and fail to capture fine-grai

Cited by 0SourcecodeScholar
2025

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

CVPR 2025highlight

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corr…

2025

Investigating Personalized Driving Behaviors in Dilemma Zones: Analysis and Prediction of Stop-or-Go Decisions

RA-L 2025

Dilemma zones at signalized intersections present a commonly occurring yet unsolved challenge in traffic safety. The onsets of yellow-light prompts varied responses from drivers: some may brake abruptly, compromising ride comfort, while others may accelerate, increasing the likelihood of red-light v

Cited by 4SourceScholar
2025

NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

ICCV 2025poster

Recent advances in multi-modal large language models (MLLMs) have demonstrated strong performance across various domains; however, their ability to comprehend driving scenes remains less proven. The complexity of driving scenarios, which includes multi-view information, poses significant challenges…

2025

On Learning Closed-Loop Probabilistic Multi-Agent Simulator

IROS 2025

The rapid iteration of autonomous vehicle (AV) deployments leads to increasing needs for building realistic and scalable multi-agent traffic simulators for efficient evaluation. Recent advances in this area focus on closed-loop simulators that enable generating diverse and interactive scenarios. Thi

Cited by 0SourceScholar
2024

LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs

CVPR 2024poster

Autonomous driving (AD) has made significant strides in recent years. However existing frameworks struggle to interpret and execute spontaneous user instructions such as "overtake the car ahead." Large Language Models (LLMs) have demonstrated impressive reasoning capabilities showing potential to br…

2023

Class Prototypes Based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos

CVPR 2023poster

The recent growth in the consumption of online media by children during early childhood necessitates data-driven tools enabling educators to filter out appropriate educational content for young learners. This paper presents an approach for detecting educational content in online videos. We focus on…

2023

Contrastive Self-Supervised Learning Leads to Higher Adversarial Susceptibility

AAAI 2023technical

Contrastive self-supervised learning (CSL) has managed to match or surpass the performance of supervised learning in image and video classification. However, it is still largely unknown if the nature of the representations induced by the two learning paradigms is similar. We investigate this under t…

Cited by 6SourcePDFScholar
2022

Personalized Car Following for Autonomous Driving with Inverse Reinforcement Learning

ICRA 2022poster

Driving automation is gradually replacing human driving maneuvers in different applications such as adaptive cruise control and lane keeping. However, contemporary driving automation applications based on expert systems or prede-fined control strategies are not in line with individual human driver's…

Cited by 59SourceScholar