← Search

Gaurav Sharma

27 accepted papers

2026

Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-Identification

AAAI 2026technical

We propose unsupervised multi-scenario (UMS) person re-identification (ReID) as a new task that expands ReID across diverse scenarios (cross-resolution, clothing change, etc.) within a single coherent framework. To tackle UMS-ReID, we introduce image-text knowledge modeling (ITKM) -- a three-stage f

Cited by 0SourcePDFScholar
2026

Predictive Display for Teleoperation Based on Vector Fields Using Lidar-Camera Fusion (Abstract Reprint)

AAAI 2026technical

Teleoperation can enable human intervention to help handle instances of failure in autonomy thus allowing for much safer deployment of autonomous vehicle technology. Successful teleoperation requires recreating the environment around the remote vehicle using camera data received over wireless commun

Cited by 0SourcePDFScholar
2025

Preserve Anything: Controllable Image Synthesis with Object Preservation

ICCV 2025poster

We introduce Preserve Anything, a novel method for con-trolled image synthesis that addresses key limitations in ob-ject preservation and semantic consistency in text-to-image(T2I) generation. Existing approaches often fail (i) to pre-serve multiple objects with fidelity, (ii) maintain semanticalign…

2024

OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

CVPR 2024poster

We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image video audio text depth point cloud time series tabular graph X-ray infrared IMU and hyperspectral. The proposed approach…

Cited by 15SourcePDFScholar
2021

Beyond Image to Depth: Improving Depth Prediction Using Echoes

CVPR 2021poster

We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utilized echoes for depth estimation. We propose an end-to-end deep learning based pi…

Cited by 48PDFcodeScholar
2021

Exploiting Local Geometry for Feature and Graph Construction for Better 3D Point Cloud Processing with Graph Neural Networks

ICRA 2021poster

We propose simple yet effective improvements in point representations and local neighborhood graph construction within the general framework of graph neural networks (GNNs) for 3D point cloud processing. As a first contribution, we propose to augment the vertex representations with important local g…

Cited by 23SourcecodeScholar
2020

Object Detection with a Unified Label Space from Multiple Datasets

ECCV 2020poster

Given multiple datasets with different label spaces, the goal of this work is to train a single object detector predicting over the union of all the label spaces. The practical benefits of such an object detector are obvious and significant---application-relevant categories can be picked and merged…

2019

Learning 2D to 3D Lifting for Object Detection in 3D for Autonomous Vehicles

IROS 2019poster

We address the problem of 3D object detection from 2D monocular images in autonomous driving scenarios. We propose to lift the 2D images to 3D representations using learned neural networks and leverage existing networks working directly on 3D data to perform 3D object detection and localization. We…

Cited by 42SourceScholar
2018

Quantification of Longitudinal Changes in Retinal Vasculature from Wide-Field Fluorescein Angiography via a Novel Registration and Change Detection Approach

ICASSP 2018accepted

Wide-field fluorescein angiography (FA) images are commonly used in ophthalmology to assess longitudinal changes in retinal vasculature, specifically, non-perfusion. Current practice relies on manual qualitative comparisons between images taken at successive clinic visits, a few months apart. Object…

Cited by 0SourceScholar
2017

AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos

CVPR 2017poster

We propose a novel method for temporally pooling frames in a video for the task of human action recognition. The method is motivated by the observation that there are only a small number of frames which, together, contain sufficient information to discriminate an action class present in a video, fro…

Cited by 201PDFcodeScholar
2017

An Empirical Evaluation of Visual Question Answering for Novel Objects

CVPR 2017poster

We study the problem of answering questions about images in the harder setting, where the test questions and corresponding images contain novel objects, which were not queried about in the training data. Such setting is inevitable in real world--owing to the heavy tailed distribution of the visual c…

Cited by 33PDFScholar
2017

In-situ calibration of accelerometers in body-worn sensors using quiescent gravity

ICASSP 2017accepted

As the cost, size and power required by sensor devices decrease, an increasing range of applications are possible. We focus on the application of tracking accelerometer data from body worn sensors over long durations in health monitoring applications. Body worn sensors must be compact for the conven…

Cited by 0SourceScholar
2017

See and listen: Score-informed association of sound tracks to players in chamber music performance videos

ICASSP 2017accepted

Both audio and visual aspects of a musical performance, especially their association, are important for expressing players' ideas and for engaging the audience. In this paper, we present a framework for combining audio and video analyses of multi-instrument chamber music performances to associate pl…

Cited by 0SourceScholar
2017

Vehicle tracking in Wide area motion imagery: A facility location motivated combinatorial approach

ICASSP 2017accepted

We propose a practical combinatorial approach for addressing the challenging task of vehicle tracking in Wide area motion imagery (WAMI) by leveraging a pixel accurate co-registered vector road-map. Specifically, guided by the co-registered road network, we obtain a sparse trellis graph linking each…

Cited by 0SourceScholar
2017

Visually informed multi-pitch analysis of string ensembles

ICASSP 2017accepted

Multi-pitch analysis of polyphonic music requires estimating concurrent pitches (estimation) and organizing them into temporal streams according to their sound sources (streaming). This is challenging for approaches based on audio alone due to the polyphonic nature of the audio signals. Video of the…

Cited by 0SourceScholar
2016

A joint approach to vector road map registration and vehicle tracking for wide area motion imagery

ICASSP 2016accepted

Modern aerial imaging platforms provide wide-area motion imagery (WAMI) at high spatial and moderate temporal resolutions making feasible a range of new applications. We consider the dual tasks of registering WAMI frames to geo-referenced vector road-maps and tracking vehicles through the progressio…

Cited by 0SourceScholar
2016

A joint learning approach for cross domain age estimation

ICASSP 2016accepted

We propose a novel joint learning method for cross domain age estimation, a domain adaptation problem. The proposed method learns a low dimensional projection along with a re-gressor, in the projection space, in a joint framework. The projection aligns the features from two different domains, i.e. s…

Cited by 0SourceScholar
2016

CP-mtML: Coupled Projection Multi-Task Metric Learning for Large Scale Face Retrieval

CVPR 2016poster

We propose a novel Coupled Projection multi-task Met- ric Learning (CP-mtML) method for large scale face re- trieval. In contrast to previous works which were limited to low dimensional features and small datasets, the proposed method scales to large datasets with high dimensional face descriptors.…

Cited by 61PDFcodeScholar
2016

Latent Embeddings for Zero-Shot Classification

CVPR 2016spotlight

We present a novel latent embedding model for learning a compatibility function between image and class embeddings, in the context of zero-shot classification. The proposed method augments the state-of-the-art bilinear compatibility model by incorporating latent variables. Instead of learning a sing…

Cited by 888PDFScholar