← Search

Bharath Hariharan

78 accepted papers

2026

Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

CVPR 2026

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settin

Cited by 0SourceScholar
2026

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

CVPR 2026

The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-grained tasks such as action segmentation, classification, captioning, and retrieval. Furthermore, these benchmarks often rely

Cited by 0SourcecodeScholar
2026

MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark

CVPR 2026

Understanding real-world videos such as movies requires integrating visual and dialogue cues. Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and, given the difficulty of evaluating free-form answers, largely resort to simple multiple choice questions. We introduce a no

Cited by 0SourceScholar
2025

Accurate Differential Operators for Hybrid Neural Fields

CVPR 2025poster

Neural fields have become widely used in various fields, from shape representation to neural rendering, and for solving partial differential equations (PDEs). With the advent of hybrid neural field representations like Instant NGP that leverage small MLPs and explicit representations, these models t…

2025

C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction

NeurIPS 2025poster

Geometric models like DUSt3R have shown great advances in understanding the geometry of a scene from pairs of photos. However, they fail when the inputs are from vastly different viewpoints (e.g., aerial vs.\ ground) or modalities (e.g., photos vs.\ abstract drawings) compared to what was observed d…

Cited by 0SourcecodeScholar
2025

DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery

CVPR 2025poster

Visual data is used in numerous different scientific workflows ranging from remote sensing to ecology. As the amount of observation data increases, the challenge is not just to make accurate predictions but also to understand the underlying mechanisms for those predictions. Good interpretation is im…

Cited by 0SourcePDFScholar
2025

FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution

ICCV 2025poster

A versatile video depth estimation model should be consistent and accurate across frames, produce high-resolution depth maps, and support real-time streaming. We propose a method, FlashDepth, that satisfies all three requirements, performing depth estimation for a 2044x1148 streaming video at 24 FPS…

2025

Generating 3D-Consistent Videos from Unposed Internet Photos

CVPR 2025poster

We address the problem of generating videos from unposed internet photos. A handful of input images serve as keyframes, and our model interpolates between them to simulate a path moving between the cameras. Given random images, a model's ability to capture underlying geometry, recognize scene identi…

Cited by 1SourcePDFScholar
2025

Learning 3D Perception from Others' Predictions

ICLR 2025poster

Accurate 3D object detection in real-world environments requires a huge amount of annotated data with high quality. Acquiring such data is tedious and expensive, and often needs repeated effort when a new sensor is adopted or when the detector is deployed in a new environment. We investigate a new s…

Cited by 1SourcePDFScholar
2025

MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing

NeurIPS 2025spotlight

Natural disasters cause devastating damage to communities and infrastructure every year. Effective disaster response is hampered by the difficulty of accessing affected areas during and after events. Remote sensing has allowed us to monitor natural disasters in a remote way. More recently there have…

Cited by 0SourceScholar
2025

Mixed Signals: A Diverse Point Cloud Dataset for Heterogeneous LiDAR V2X Collaboration

ICCV 2025poster

Vehicle-to-everything (V2X) collaborative perception has emerged as a promising solution to address the limitations of single-vehicle perception systems. However, existing V2X datasets are limited in scope, diversity, and quality. To address these gaps, we present Mixed Signals, a comprehensive V2X…

Cited by 0SourcePDFScholar
2025

Scale-aware Recognition in Satellite Images under Resource Constraints

ICLR 2025poster

Recognition of features in satellite imagery (forests, swimming pools, etc.) depends strongly on the spatial scale of the concept and therefore the resolution of the images. This poses two challenges: Which resolution is best suited for recognizing a given concept, and where and when should the cost…

Cited by 0SourcePDFScholar
2025

Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene

CVPR 2025poster

Self-driving cars relying solely on ego-centric perception face limitations in sensing, often failing to detect occluded, faraway objects. Collaborative autonomous driving (CAV) seems like a promising direction, but collecting data for development is non-trivial. It requires placing multiple sensor-…

Cited by 0SourcePDFScholar
2024

AllClear: A Comprehensive Dataset and Benchmark for Cloud Removal in Satellite Imagery

NeurIPS 2024poster

Clouds in satellite imagery pose a significant challenge for downstream applications. A major challenge in current cloud removal research is the absence of a comprehensive benchmark and a sufficiently large and diverse training dataset. To address this problem, we introduce the largest public datase…

2024

Better Monocular 3D Detectors with LiDAR from the Past

ICRA 2024poster

Accurate 3D object detection is crucial to autonomous driving. Though LiDAR-based detectors have achieved impressive performance, the high cost of LiDAR sensors precludes their widespread adoption in affordable vehicles. Camera-based detectors are cheaper alternatives but often suffer inferior perfo…

Cited by 1SourcecodeScholar
2024

Counter-Current Learning: A Biologically Plausible Dual Network Approach for Deep Learning

NeurIPS 2024poster

Despite its widespread use in neural networks, error backpropagation has faced criticism for its lack of biological plausibility, suffering from issues such as the backward locking problem and the weight transport problem. These limitations have motivated researchers to explore more biologically pl…

2024

DiffuBox: Refining 3D Object Detection with Point Diffusion

NeurIPS 2024poster

Ensuring robust 3D object detection and localization is crucial for many applications in robotics and autonomous driving. Recent models, however, face difficulties in maintaining high performance when applied to domains with differing sensor setups or geographic locations, often resulting in poor lo…

2024

MegaScenes: Scene-Level View Synthesis at Scale

ECCV 2024poster

"Scene-level novel view synthesis (NVS) is fundamental to many vision and graphics applications. Recently, pose-conditioned diffusion models have led to significant progress by extracting 3D information from 2D foundation models, but these methods are limited by the lack of scene-level training data…

2024

Pre-training LiDAR-based 3D Object Detectors through Colorization

ICLR 2024poster

Accurate 3D object detection and understanding for self-driving cars heavily relies on LiDAR point clouds, necessitating large amounts of labeled data to train. In this work, we introduce an innovative pre-training approach, Grounded Point Colorization (GPC), to bridge the gap between data and label…

2024

Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment

ICLR 2024poster

We introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for connecting remote-sensing images and language. Specifically, we train an image enco…

Cited by 50SourcePDFScholar
2023

Change-Aware Sampling and Contrastive Learning for Satellite Images

CVPR 2023poster

Automatic remote sensing tools can help inform many large-scale challenges such as disaster management, climate change, etc. While a vast amount of spatio-temporal satellite image data is readily available, most of it remains unlabelled. Without labels, this data is not very useful for supervised le…

2023

Doppelgangers: Learning to Disambiguate Images of Similar Structures

ICCV 2023oral

We consider the visual disambiguation task of determining whether a pair of visually similar images depict the same or distinct 3D surfaces (e.g., the same or opposite sides of a symmetric building). Illusory image matches, where two images observe distinct but visually similar 3D surfaces, can be c…

Cited by 37PDFcodeScholar
2023

Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical Scenes

NeurIPS 2023poster

Unsupervised monocular depth estimation techniques have demonstrated encouraging results but typically assume that the scene is static. These techniques suffer when trained on dynamical scenes, where apparent object motion can equally be explained by hypothesizing the object's independent motion, or…

2023

Emergent Correspondence from Image Diffusion

NeurIPS 2023poster

Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We propose a simple strategy to extract this implicit knowledge out of diffusion networks as image featur…

2023

Image-to-Image Translation for Autonomous Driving from Coarsely-Aligned Image Pairs

ICRA 2023poster

A self-driving car must be able to reliably handle adverse weather conditions (e.g., snowy) to operate safely. In this paper, we investigate the idea of turning sensor inputs (i.e., images) captured in an adverse condition into a benign one (i.e., sunny), upon which the downstream tasks (e.g., seman…

Cited by 18SourceScholar
2023

Reward Finetuning for Faster and More Accurate Unsupervised Object Discovery

NeurIPS 2023poster

Recent advances in machine learning have shown that Reinforcement Learning from Human Feedback (RLHF) can improve machine learning models and align them with human preferences. Although very successful for Large Language Models (LLMs), these advancements have not had a comparable impact in research…

2023

Tracking Everything Everywhere All at Once

ICCV 2023oral

We present a new test-time optimization method for estimating dense and long-range motion from a video sequence. Prior optical flow or particle video tracking algorithms typically operate within limited temporal windows, struggling to track through occlusions and maintain global consistency of estim…

Cited by 170PDFcodeScholar
2022

Change Event Dataset for Discovery from Spatio-temporal Remote Sensing Imagery

NeurIPS 2022accept

Satellite imagery is increasingly available, high resolution, and temporally detailed. Changes in spatio-temporal datasets such as satellite images are particularly interesting as they reveal the many events and forces that shape our world. However, finding such interesting and meaningful change e…

Cited by 17SourcePDFScholar
2022

Exploiting Playbacks in Unsupervised Domain Adaptation for 3D Object Detection in Self-Driving Cars

ICRA 2022poster

Self-driving cars must detect other traffic participants like vehicles and pedestrians in 3D in order to plan safe routes and avoid collisions. State-of-the-art 3D object detectors, based on deep learning, have shown promising accuracy but are prone to over-fit domain idiosyncrasies, making them fai…

Cited by 25SourceScholar
2022

Hindsight is 20/20: Leveraging Past Traversals to Aid 3D Perception

ICLR 2022poster

Self-driving cars must detect vehicles, pedestrians, and other traffic participants accurately to operate safely. Small, far-away, or highly occluded objects are particularly challenging because there is limited information in the LiDAR point clouds for detecting them. To address this challenge, we l…

2022

Ithaca365: Dataset and Driving Perception Under Repeated and Challenging Weather Conditions

CVPR 2022poster

Advances in perception for self-driving cars have accelerated in recent years due to the availability of large-scale datasets, typically collected at specific locations and under nice weather conditions. Yet, to achieve the high safety requirement, these perceptual systems must operate robustly unde…

Cited by 53PDFScholar
2022

Learning To Detect Mobile Objects From LiDAR Scans Without Labels

CVPR 2022poster

Current 3D object detectors for autonomous driving are almost entirely trained on human-annotated data. Although of high quality, the generation of such data is laborious and costly, restricting them to a few specific locations and object types. This paper proposes an alternative approach entirely b…

Cited by 51PDFcodeScholar
2022

Polynomial Neural Fields for Subband Decomposition and Manipulation

NeurIPS 2022accept

Neural fields have emerged as a new paradigm for representing signals, thanks to their ability to do it compactly while being easy to optimize. In most applications, however, neural fields are treated like a black box, which precludes many signal manipulation tasks. In this paper, we propose a new c…

2022

Unsupervised Adaptation from Repeated Traversals for Autonomous Driving

NeurIPS 2022accept

For a self-driving car to operate reliably, its perceptual system must generalize to the end-user's environment --- ideally without additional annotation efforts. One potential solution is to leverage unlabeled data (e.g., unlabeled LiDAR point clouds) collected from the end-users' environments (i.e…

2022

Unsupervised Semantic Segmentation by Distilling Feature Correspondences

ICLR 2022poster

Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clus…

2022

Visual Prompt Tuning

ECCV 2022poster

"The current modus operandi in adapting pre-trained models involves updating all the backbone parameters, i.e. full fine-tuning. This paper introduces Visual Prompt Tuning (VPT) as an efficient and effective alternative to full fine-tuning for large-scale Transformer models in vision. Taking inspira…

2021

Extreme Rotation Estimation Using Dense Correlation Volumes

CVPR 2021poster

We present a technique for estimating the relative 3D rotation of an RGB image pair in an extreme setting, where the images have little or no overlap. We observe that, even when images do not overlap, there may be rich hidden cues as to their geometric relationship, such as light source directions,…

Cited by 48PDFcodeScholar
2021

Few-Shot Classification With Feature Map Reconstruction Networks

CVPR 2021poster

In this paper we reformulate few-shot classification as a reconstruction problem in latent space. The ability of the network to reconstruct a query feature map from support features of a given class predicts membership of the query in that class. We introduce a novel mechanism for few-shot classific…

Cited by 357PDFcodeScholar
2021

PiCIE: Unsupervised Semantic Segmentation Using Invariance and Equivariance in Clustering

CVPR 2021poster

We present a new framework for semantic segmentation without annotations via clustering. Off-the-shelf clustering methods are limited to curated, single-label, and object-centric images yet real-world data are dominantly uncurated, multi-label, and scene-centric. We extend clustering from images to…

Cited by 238PDFcodeScholar
2021

Self-training For Few-shot Transfer Across Extreme Task Differences

ICLR 2021oral

Most few-shot learning techniques are pre-trained on a large, labeled “base dataset”. In problem domains where such large labeled datasets are not available for pre-training (e.g., X-ray, satellite images), one must resort to pre-training in a different “source” problem domain (e.g., ImageNet), whic…

2021

Stay Positive: Non-Negative Image Synthesis for Augmented Reality

CVPR 2021poster

In applications such as optical see-through and projector augmented reality, producing images amounts to solving non-negative image generation, where one can only add light to an existing image. Most image generation methods, however, are ill-suited to this problem setting, as they make the assumpti…

Cited by 7PDFcodeScholar
2020

End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection

CVPR 2020poster

Reliable and accurate 3D object detection is a necessity for safe autonomous driving. Although LiDAR sensors can provide accurate 3D point cloud estimates of the environment, they are also prohibitively expensive for many settings. Recently, the introduction of pseudo-LiDAR (PL) has led to a drastic…

Cited by 262PDFcodeScholar
2020

Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset

ECCV 2020poster

Segmentation, and an Attribute Localization Dataset","In this work, we focus on the task of instance segmentation with attribute localization. This unifies instance segmentation (detect and segment each object instance) and visual categorization of fine-grained attributes (classify one or multiple a…

2020

Learning Feature Descriptors using Camera Pose Supervision

ECCV 2020poster

Recent research on learned visual descriptors has shown promising improvements in correspondence estimation, a key component of many 3D vision tasks. However, existing descriptor learning frameworks typically require ground-truth correspondences between feature points for training, which are challen…

Cited by 201SourcePDFScholar
2020

Learning Gradient Fields for Shape Generation

ECCV 2020poster

In this work, we propose a novel technique to generate shapes from point cloud data. A point cloud can be viewed as samples from a distribution of 3D points whose density is concentrated near the surface of the shape. Point cloud generation thus amounts to moving randomly sampled points to high-dens…

2020

Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous Driving

ICLR 2020poster

Detecting objects such as cars and pedestrians in 3D plays an indispensable role in autonomous driving. Existing approaches largely rely on expensive LiDAR sensors for accurate depth information. While recently pseudo-LiDAR has been introduced as a promising alternative, at a much lower cost based s…

Cited by 515SourcecodeScholar
2020

Revisiting Pose-Normalization for Fine-Grained Few-Shot Recognition

CVPR 2020poster

Few-shot, fine-grained classification requires a model to learn subtle, fine-grained distinctions between different classes (e.g., birds) based on a few images alone. This requires a remarkable degree of invariance to pose, articulation and background. A solution is to use pose-normalized representa…

Cited by 69PDFcodeScholar
2020

Train in Germany, Test in the USA: Making 3D Object Detectors Generalize

CVPR 2020poster

In the domain of autonomous driving, deep learning has substantially improved the 3D object detection accuracy for LiDAR and stereo camera data alike. While deep networks are great at generalization, they are also notorious to overfit to all kinds of spurious artifacts, such as brightness, car sizes…

Cited by 215PDFcodeScholar
2020

Wasserstein Distances for Stereo Disparity Estimation

NeurIPS 2020spotlight

Existing approaches to depth or disparity estimation output a distribution over a set of pre-defined discrete values. This leads to inaccurate results when the true depth or disparity does not match any of these values. The fact that this distribution is usually learned indirectly through a regressi…

2019

LDLS: 3-D Object Segmentation Through Label Diffusion From 2-D Images

RA-L 2019

Object segmentation in three-dimensional (3-D) point clouds is a critical task for robots capable of 3-D perception. Despite the impressive performance of deep learning-based approaches on object segmentation in 2-D images, deep learning has not been applied nearly as successfully for 3-D point clou

Cited by 39SourcecodeScholar
2019

PointFlow: 3D Point Cloud Generation With Continuous Normalizing Flows

ICCV 2019oral

As 3D point clouds become the representation of choice for multiple vision and graphics applications, the ability to synthesize or reconstruct high-resolution, high-fidelity point clouds becomes crucial. Despite the recent success of deep learning models in discriminative tasks of point clouds, gene…

Cited by 798PDFcodeScholar
2019

Pseudo-LiDAR From Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving

CVPR 2019poster

3D object detection is an essential task in autonomous driving. Recent techniques excel with highly accurate detection rates, provided the 3D input data is obtained from precise but expensive LiDAR technology. Approaches based on cheaper monocular or stereo imagery data have, until now, resulted in…

Cited by 1347PDFcodeScholar
2018

Learning Single-View 3D Reconstruction with Limited Pose Supervision

ECCV 2018poster

It is expensive to label images with 3D structure or precise camera pose. Yet, this is precisely the kind of annotation required to train single-view 3D reconstruction models. In contrast, unlabeled images or images with just category labels are easy to acquire, but few current models can use this w…

2018

Low-Shot Learning With Large-Scale Diffusion

CVPR 2018poster

This paper considers the problem of inferring image labels from images when only a few annotated examples are available at training time. This setup is often referred to as low-shot learning, where a standard approach is to re-train the last few layers of a convolutional neural network learned on s…

2018

Resource Aware Person Re-Identification Across Multiple Resolutions

CVPR 2018poster

Not all people are equally easy to identify: color statistics might be enough for some cases while others might require careful reasoning about high- and low-level details. However, prevailing person re-identification(re-ID) methods use one-size-fits-all high-level embeddings from deep convolutional…

2017

CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

CVPR 2017poster

When building artificial intelligence systems that can reason and answer questions about visual data, we need diagnostic tests to analyze our progress and discover short- comings. Existing benchmarks for visual question answer- ing can help, but have strong biases that models can exploit to correctl…

Cited by 2819PDFScholar
2017

Feature Pyramid Networks for Object Detection

CVPR 2017poster

Feature pyramids are a basic component in recognition systems for detecting objects at different scales. But pyramid representations have been avoided in recent object detectors that are based on deep convolutional networks, partially because they are slow to compute and memory intensive. In this pa…

Cited by 33044PDFcodeScholar
2017

Inferring and Executing Programs for Visual Reasoning

ICCV 2017oral

Existing methods for visual reasoning attempt to directly map inputs to outputs using black-box architectures without explicitly modeling the underlying reasoning processes. As a result, these black-box models often learn to exploit biases in the data rather than learning to perform visual reasoning…

Cited by 677PDFcodeScholar
2017

Learning Features by Watching Objects Move

CVPR 2017poster

This paper presents a novel yet intuitive approach to unsupervised feature learning. Inspired by the human visual system, we explore whether low-level motion-based grouping cues can be used to learn an effective visual representation. Specifically, we use unsupervised motion-based segmentation on vi…

Cited by 640PDFcodeScholar
2015

Hypercolumns for Object Segmentation and Fine-Grained Localization

CVPR 2015poster

Recognition algorithms based on convolutional networks (CNNs) typically use the output of the last layer as feature representation. However, the information in this layer may be too coarse to allow precise localization. On the contrary, earlier layers may be precise in localization but will not capt…

Cited by 2008SourcePDFScholar