← Search

Vineet Gandhi

23 accepted papers

2026

Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach

CVPR 2026

Contrastive vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition yet remain vulnerable to spurious correlations, particularly background over-reliance. We introduce Cluster-based Concept Importance (CCI), a novel interpretability method that uses CLIP's own patch embedding

Cited by 0SourceScholar
2025

Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset

ICASSP 2025accepted

Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well across different speakers. To address this issue, we focus on learning phoneme-leve…

Cited by 0SourceScholar
2025

IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs

NAACL 2025short

Recent evaluations of LLMs on coreference resolution have revealed that traditional output formats and evaluation metrics do not fully capture the models’ referential understanding. To address this, we introduce IdentifyMe, a new benchmark for mention resolution presented in a multiple-choice questi…

2025

MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI

ICASSP 2025accepted

Previous real-time MRI (rtMRI)-based speech synthesis models depend heavily on noisy ground-truth speech. Applying loss directly over ground truth mel-spectrograms entangles speech content with MRI noise, resulting in poor intelligibility. We introduce a novel approach that adapts the multi-modal se…

Cited by 0SourceScholar
2025

Minimalistic Video Saliency Prediction via Efficient Decoder & Spatio Temporal Action Cues

ICASSP 2025accepted

This paper introduces ViNet-S, a 36MB model based on the ViNet architecture with a U-Net design, featuring a lightweight decoder that significantly reduces model size and parameters without compromising performance. Additionally, ViNet-A (148MB) incorporates spatio-temporal action localization (STAL…

Cited by 0SourceScholar
2025

Prompt-to-Correct: Automated Test-Time Pronunciation Correction with Voice Prompts

ICASSP 2025accepted

Pronunciation correction is crucial for Text-to-Speech (TTS) systems in production. Traditional methods, which rely on phoneme sequence manipulation, are often cumbersome and error-prone. To address this, we propose Prompt-to-Correct, an editing-based methodology for pronunciation correction in TTS…

Cited by 0SourceScholar
2025

TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction

CVPR 2025highlight

We consider the problem of single-source domain generalization. Existing methods typically rely on extensive augmentations to synthetically cover diverse domains during training. However, they struggle with semantic shifts (e.g., background and viewpoint changes), as they often learn global features…

Cited by 2SourcePDFScholar
2025

VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

CVPR 2025poster

A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision/video models and a move towards long-video understanding. While exciting, we take a step back and ask: are today's models go…

2024

Major Entity Identification: A Generalizable Alternative to Coreference Resolution

EMNLP 2024main

The limited generalization of coreference resolution (CR) models has been a major bottleneck in the task’s broad application. Prior work has identified annotation differences, especially for mention detection, as one of the main reasons for the generalization gap and proposed using additional annota…

2023

Ground then Navigate: Language-guided Navigation in Dynamic Scenes

ICRA 2023poster

We investigate the Vision-and-Language Navigation (VLN) problem in the context of autonomous driving in outdoor settings. We solve the problem by explicitly grounding the navigable regions corresponding to the textual command. At each timestamp, the model predicts a segmentation mask corresponding t…

Cited by 27SourcecodeScholar
2023

Test-Time Amendment with a Coarse Classifier for Fine-Grained Classification

NeurIPS 2023poster

We investigate the problem of reducing mistake severity for fine-grained classification. Fine-grained classification can be challenging, mainly due to the requirement of knowledge or domain expertise for accurate annotation. However, humans are particularly adept at performing coarse classification…

2022

Empathic Machines: Using Intermediate Features as Levers to Emulate Emotions in Text-To-Speech Systems

NAACL 2022long

We present a method to control the emotional prosody of Text to Speech (TTS) systems by using phoneme-level intermediate features (pitch, energy, and duration) as levers. As a key idea, we propose Differential Scaling (DS) to disentangle features relating to affective prosody from those arising due…

2021

Grounding Linguistic Commands to Navigable Regions

IROS 2021poster

Humans have a natural ability to effortlessly comprehend linguistic commands such as “park next to the yellow sedan” and instinctively know which region of the road the vehicle should navigate. Extending this ability to autonomous vehicles is the next step towards creating fully autonomous agents th…

Cited by 14SourcecodeScholar
2021

No Cost Likelihood Manipulation at Test Time for Making Better Mistakes in Deep Networks

ICLR 2021poster

There has been increasing interest in building deep hierarchy-aware classifiers that aim to quantify and reduce the severity of mistakes, and not just reduce the number of errors. The idea is to exploit the label hierarchy (e.g., the WordNet ontology) and consider graph distances as a proxy for mist…

2021

ViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction

IROS 2021poster

We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the decoder infers a saliency map via trilinear interpolation and 3D convolutions,…

Cited by 100SourcecodeScholar
2020

LiDAR guided Small obstacle Segmentation

IROS 2020poster

Detecting small obstacles on the road is critical for autonomous driving. In this paper, we present a method to reliably detect such obstacles through a multi-modal framework of sparse LiDAR(VLP-16) and Monocular vision. LiDAR is employed to provide additional context in the form of confidence maps…

Cited by 33SourcecodeScholar
2019

Nose, Eyes and Ears: Head Pose Estimation by Locating Facial Keypoints

ICASSP 2019accepted

Monocular head pose estimation requires learning a model that computes the intrinsic Euler angles for pose (yaw, pitch, roll) from an input image of human face. Annotating ground truth head pose angles for images in the wild is difficult and requires ad-hoc fitting procedures (which provides only co…

Cited by 0SourceScholar
2019

Talk to the Vehicle: Language Conditioned Autonomous Navigation of Self Driving Cars

IROS 2019poster

We propose a novel pipeline that blends encodings from natural language and 3D semantic maps obtained from visual imagery to generate local trajectories that are executed by a low-level controller. The pipeline precludes the need for a prior registered map through a local waypoint generator neural n…

Cited by 30SourceScholar
2018

MergeNet: A Deep Net Architecture for Small Obstacle Discovery

ICRA 2018poster

We present here, a novel network architecture called MergeNet for discovering small obstacles for on-road scenes in the context of autonomous driving. The basis of the architecture rests on the central consideration of training with less amount of data since the physical setup and the annotation pro…

Cited by 29SourceScholar