← Search

Deva Ramanan

129 accepted papers

2026

Any4D: Unified Feed-Forward Metric 4D Reconstruction

CVPR 2026

We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N frames, in contrast to prior work that typically focuses on either 2-view dense scene flow or sparse 3D point tracking.

Cited by 0SourcecodeScholar
2026

Building a Precise Video Language with Human-AI Oversight

CVPR 2026

Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable oversight that enable precise video captioning. First, we define a structured specification for describing subjects, scene

Cited by 0SourcecodeScholar
2026

Contact-guided Real2Sim from Monocular Video with Planar Scene Primitives

ICLR 2026poster

We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human--scene reconstruction relies on data-driven priors and joint optimization with no physics in the loop, or recovers noisy geometry with artifacts that cause motion-tr…

Cited by 0SourcecodeScholar
2026

RF-DETR: Neural Architecture Search for Real-Time Detection Transformers

ICLR 2026poster

Open-vocabulary detectors achieve impressive performance on COCO, but often fail to generalize to real-world datasets with out-of-distribution classes not typically found in their pre-training. Rather than simply fine-tuning a heavy-weight vision-language model (VLM) for new domains, we introduce RF…

Cited by 0SourcecodeScholar
2025

AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis

CVPR 2025poster

We explore the task of geometric reconstruction of images captured from a mixture of ground and aerial views. Current state-of-the-art learning-based approaches fail to handle the extreme viewpoint variation between aerial-ground image pairs. Our hypothesis is that the lack of high-quality, co-regis…

Cited by 1SourcePDFScholar
2025

BETTY Dataset: A Multi-Modal Dataset for Full-Stack Autonomy

ICRA 2025

We present the BETTY dataset, a large-scale, multi-modal dataset collected on several autonomous racing vehicles, targeting supervised and self-supervised state estimation, dynamics modeling, motion forecasting, perception, and more. Existing large-scale datasets, especially autonomous vehicle datas

Cited by 1SourcecodeScholar
2025

DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion

CVPR 2025poster

Current Structure-from-Motion (SfM) methods typically follow a two-stage pipeline, combining learned or geometric pairwise reasoning with a subsequent global optimization step. In contrast, we propose a data-driven multi-view reasoning approach that directly infers 3D scene geometry and camera poses…

2025

Efficient Autoregressive Shape Generation via Octree-Based Adaptive Tokenization

ICCV 2025poster

Many 3D generative models rely on variational autoencoders (VAEs) to learn compact shape representations. However, existing methods encode all shapes into a fixed-size token, disregarding the inherent variations in scale and complexity across 3D data. This leads to inefficient latent representations…

Cited by 0SourcePDFScholar
2025

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

ICCV 2025poster

Generative Large Multimodal Models (LMMs) like LLaVA and Qwen-VL excel at a wide variety of vision-language (VL) tasks. Despite strong performance, LMMs' generative outputs are not specialized for vision-language classification tasks (i.e., tasks with vision-language inputs and discrete labels) such…

Cited by 0SourcePDFScholar
2025

Generating Physically Stable and Buildable Brick Structures from Text

ICCV 2025poster

We introduce BrickGPT, the first approach for generating physically stable interconnecting brick assembly models from text prompts. To achieve this, we construct a large-scale, physically stable dataset of brick structures, along with their associated captions, and train an autoregressive large lang…

2025

InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning

ACL 2025long

Large multimodal foundation models, particularly in the domains of language and vision, have significantly advanced various tasks, including robotics, autonomous driving, information retrieval, and grounding. However, many of these models perceive objects as indivisible, overlooking the components t…

2025

MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion

ICCV 2025poster

We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panoptic Studio) - such multi-view setups are prohibitively expensive to build and cannot capture diverse scenes in-the-wild.…

Cited by 0SourcePDFScholar
2025

Neural Eulerian Scene Flow Fields

ICLR 2025poster

We reframe scene flow as the task of estimating a continuous space-time ordinary differential equation (ODE) that describes motion for an entire observation sequence, represented with a neural prior. Our method, EulerFlow, optimizes this neural prior estimate against several multi-observation recons…

Cited by 1SourcePDFScholar
2025

ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models

ICCV 2025poster

Recent Large Vision-Language Models (LVLMs) have introduced a new paradigm for understanding and reasoning about image input through textual responses. Although they have achieved remarkable performance across a range of multi-modal tasks, they face the persistent challenge of hallucination, which i…

2025

RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion

NeurIPS 2025poster

3D shape completion has broad applications in robotics, digital twin reconstruction, and extended reality (XR). Although recent advances in 3D object and scene completion have achieved impressive results, existing methods lack 3D consistency, are computationally expensive, and struggle to capture sh…

Cited by 0SourceScholar
2025

Reanimating Images using Neural Representations of Dynamic Stimuli

CVPR 2025poster

While computer vision models have made incredible strides in static image recognition, they still do not match human performance in tasks that require the understanding of complex, dynamic motion. This is notably true for real-world scenarios where embodied agents face complex and motion-rich enviro…

2025

Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos

NeurIPS 2025poster

We explore novel-view synthesis for dynamic scenes from monocular videos. Prior approaches rely on costly test-time optimization of 4D representations or do not preserve scene geometry when trained in a feed-forward manner. Our approach is based on three key insights: (1) covisible pixels (that are…

Cited by 0SourceScholar
2025

Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models

NeurIPS 2025poster

Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of-the-art models still struggle to generalize to out-of-distribution classes, tasks and imaging modalities not typically…

Cited by 0SourcecodeScholar
2025

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

ICLR 2025poster

While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text responses that do not align with the given visual input, which restricts their practical applicability in real-world scenarios. In this work, insp…

2025

Towards Foundational Models for Single-Chip Radar

ICCV 2025poster

mmWave radars are compact, inexpensive, and durable sensors that are robust to occlusions and work regardless of environmental conditions, such as weather and darkness. However, this comes at the cost of poor angular resolution, especially for inexpensive single-chip radars, which are typically used…

Cited by 0SourcePDFScholar
2025

Towards Understanding Camera Motions in Any Video

NeurIPS 2025spotlight

We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous multi-stage quality control process. One of our core contributions is a taxonomy o…

Cited by 0SourceScholar
2025

UFM: A Simple Path towards Unified Dense Correspondence with Flow

NeurIPS 2025poster

Dense image correspondence is central to many applications, such as visual odometry, 3D reconstruction, object association, and re-identification. Historically, dense correspondence has been tackled separately for wide-baseline scenarios and optical flow estimation, despite the common goal of matchi…

Cited by 0SourceScholar
2024

Better Call SAL: Towards Learning to Segment Anything in Lidar

ECCV 2024poster

"We propose the (Segment Anything in Lidar) method consisting of a text-promptable zero-shot model for segmenting and classifying any object in Lidar, and a pseudo-labeling engine that facilitates model training without manual supervision. While the established paradigm for (LPS) relies on manual su…

2024

Cameras as Rays: Pose Estimation via Ray Diffusion

ICLR 2024oral

Estimating camera poses is a fundamental task for 3D reconstruction and remains challenging given sparsely sampled views (<10). In contrast to existing approaches that pursue top-down prediction of global parametrizations of camera extrinsics, we propose a distributed representation of camera pose t…

Cited by 60SourcePDFScholar
2024

Evaluating Text-to-Visual Generation with Image-to-Text Generation

ECCV 2024poster

"Despite significant progress in generative AI, comprehensive evaluation remains challenging because of the lack of effective metrics and standardized benchmarks. For instance, the widely-used CLIPScore measures the alignment between a (generated) image and text prompt, but it fails to produce relia…

2024

FlashTex: Fast Relightable Mesh Texturing with LightControlNet

ECCV 2024oral

"Manually creating textures for 3D meshes is time-consuming, even for expert visual content creators. We propose a fast approach for automatically texturing an input 3D mesh based on a user-provided text prompt. Importantly, our approach disentangles lighting from surface material/reflectance in the…

Cited by 28SourcePDFScholar
2024

HybridNeRF: Efficient Neural Rendering via Adaptive Volumetric Surfaces

CVPR 2024highlight

Neural radiance fields provide state-of-the-art view synthesis quality but tend to be slow to render. One reason is that they make use of volume rendering thus requiring many samples (and model queries) per ray at render time. Although this representation is flexible and easy to optimize most real-w…

Cited by 21SourcePDFScholar
2024

Language Models as Black-Box Optimizers for Vision-Language Models

CVPR 2024poster

Vision-language models (VLMs) pre-trained on web-scale datasets have demonstrated remarkable capabilities on downstream tasks when fine-tuned with minimal data. However many VLMs rely on proprietary data and are not open-source which restricts the use of white-box approaches for fine-tuning. As such…

2024

NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

NeurIPS 2024poster

Vision-language models (VLMs) have made significant progress in recent visual-question-answering (VQA) benchmarks that evaluate complex visio-linguistic reasoning. However, are these models truly effective? In this work, we show that VLMs still struggle with natural images and questions that humans…

Cited by 19SourcePDFScholar
2024

Revisiting Few-Shot Object Detection with Vision-Language Models

NeurIPS 2024poster

The era of vision-language models (VLMs) trained on web-scale datasets challenges conventional formulations of “open-world" perception. In this work, we revisit the task of few-shot object detection (FSOD) in the context of recent foundational VLMs. First, we point out that zero-shot predictions fro…

2024

Revisiting the Role of Language Priors in Vision-Language Models

ICML 2024poster

Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning. We study $\textit{generative VLMs}$ that are trained for next-word generation given an image. We explore their zero-shot perform…

2024

Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection

CoRL 2024poster

State-of-the-art 3D object detectors are often trained on massive labeled datasets. However, annotating 3D bounding boxes remains prohibitively expensive and time-consuming, particularly for LiDAR. Instead, recent works demonstrate that self-supervised pre-training with unlabeled data can improve de…

Cited by 0SourcecodeScholar
2024

SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM

CVPR 2024poster

Dense simultaneous localization and mapping (SLAM) is crucial for robotics and augmented reality applications. However current methods are often hampered by the non-volumetric or implicit way they represent a scene. This work introduces SplaTAM an approach that for the first time leverages explicit…

2024

The Neglected Tails in Vision-Language Models

CVPR 2024poster

Vision-language models (VLMs) excel in zero-shot recognition but their performance varies greatly across different visual concepts. For example although CLIP achieves impressive accuracy on ImageNet (60-80%) its performance drops below 10% for more than ten concepts like night snake presumably due t…

Cited by 44SourcePDFScholar
2024

ZeroFlow: Scalable Scene Flow via Distillation

ICLR 2024poster

Scene flow estimation is the task of describing the 3D motion field between temporally successive point clouds. State-of-the-art methods use strong priors and test-time optimization techniques, but require on the order of tens of seconds to process full-size point clouds, making them unusable as com…

2023

Joint Metrics Matter: A Better Standard for Trajectory Forecasting

ICCV 2023poster

Multi-modal trajectory forecasting methods commonly evaluate using single-agent metrics (marginal metrics), such as minimum Average Displacement Error (ADE) and Final Displacement Error (FDE), which fail to capture joint performance of multiple interacting agents. Only focusing on marginal metrics c…

Cited by 13PDFcodeScholar
2023

Learning Lightweight Object Detectors via Multi-Teacher Progressive Distillation

ICML 2023poster

Resource-constrained perception systems such as edge computing and vision-for-robotics require vision models to be both accurate and lightweight in computation and memory usage. While knowledge distillation is a proven strategy to enhance the performance of lightweight classification models, its app…

2023

Lidar Panoptic Segmentation and Tracking without Bells and Whistles

IROS 2023poster

State-of-the-art lidar panoptic segmentation (LPS) methods follow “bottom-up” segmentation-centric fashion wherein they build upon semantic segmentation networks by utilizing clustering to obtain object instances. In this paper, we re-think this approach and propose a surprisingly simple yet effecti…

Cited by 8SourcecodeScholar
2023

Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning With Multimodal Models

CVPR 2023poster

The ability to quickly learn a new task with minimal instruction - known as few-shot learning - is a central aspect of intelligent agents. Classical few-shot benchmarks make use of few-shot samples from a single modality, but such samples may not be sufficient to characterize an entire concept class…

2023

PPR: Physically Plausible Reconstruction from Monocular Videos

ICCV 2023oral

Given monocular videos, we build 3D models of articulated objects and environments whose 3D configurations satisfy dynamics and contact constraints. At its core, our method leverages differentiable physics simulation to aid visual reconstructions. We couple differentiable physics simulation with dif…

Cited by 31PDFcodeScholar
2023

Pix2map: Cross-Modal Retrieval for Inferring Street Maps From Images

CVPR 2023poster

Self-driving vehicles rely on urban street maps for autonomous navigation. In this paper, we introduce Pix2Map, a method for inferring urban street map topology directly from ego-view images, as needed to continually update and expand existing maps. This is a challenging task, as we need to infer a…

Cited by 10SourcePDFScholar
2023

Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting

CVPR 2023poster

Predicting how the world can evolve in the future is crucial for motion planning in autonomous systems. Classical methods are limited because they rely on costly human annotations in the form of semantic class labels, bounding boxes, and tracks or HD maps of cities to plan their motion -- and thus a…

2023

PyNeRF: Pyramidal Neural Radiance Fields

NeurIPS 2023poster

Neural Radiance Fields (NeRFs) can be dramatically accelerated by spatial grid representations. However, they do not explicitly reason about scale and so introduce aliasing artifacts when reconstructing scenes captured at different camera distances. Mip-NeRF and its extensions propose scale-aware re…

2023

Reconstructing Animatable Categories From Videos

CVPR 2023poster

Building animatable 3D models is challenging due to the need for 3D scans, laborious registration, and manual rigging. Recently, differentiable rendering provides a pathway to obtain high-quality 3D models from monocular videos, but these are limited to rigid categories or single instances. We prese…

2023

SLoMo: A General System for Legged Robot Motion Imitation From Casual Videos

RA-L 2023

We present SLoMo: a first-of-its-kind framework for transferring skilled motions from casually captured “in-the-wild” video footage of humans and animals to legged robots. SLoMo works in three stages: 1) synthesize a physically plausible reconstructed key-point trajectory from monocular videos; 2) o

Cited by 29SourcecodeScholar
2023

Soft Augmentation for Image Classification

CVPR 2023poster

Modern neural networks are over-parameterized and thus rely on strong regularization such as data augmentation and weight decay to reduce overfitting and improve generalization. The dominant form of data augmentation applies invariant transforms, where the learning target of a sample is invariant to…

2023

TarViS: A Unified Approach for Target-Based Video Segmentation

CVPR 2023highlight

The general domain of video segmentation is currently fragmented into different tasks spanning multiple benchmarks. Despite rapid progress in the state-of-the-art, current methods are overwhelmingly task-specific and cannot conceptually generalize to other tasks. Inspired by recent approaches with m…

2023

Total-Recon: Deformable Scene Reconstruction for Embodied View Synthesis

ICCV 2023poster

We explore the task of embodied view synthesis from monocular videos of deformable scenes. Given a minute-long RGBD video of people interacting with their pets, we render the scene from novel camera trajectories derived from the in-scene motion of actors: (1) egocentric cameras that simulate the poi…

Cited by 20PDFcodeScholar
2022

BANMo: Building Animatable 3D Neural Models From Many Casual Videos

CVPR 2022oral

Prior work for articulated 3D shape reconstruction often relies on specialized multi-view and depth sensors or pre-built deformable 3D models. Such methods do not scale to diverse sets of objects in the wild. We present a method that requires neither of them. It builds high-fidelity, articulated 3D…

Cited by 202PDFcodeScholar
2022

Continual Learning with Evolving Class Ontologies

NeurIPS 2022accept

Lifelong learners must recognize concept vocabularies that evolve over time. A common yet underexplored scenario is learning with class labels that continually refine/expand old classes. For example, humans learn to recognize ${\tt dog}$ before dog breeds. In practical settings, dataset ${\it versio…

Cited by 12SourcePDFScholar
2022

Depth-Supervised NeRF: Fewer Views and Faster Training for Free

CVPR 2022poster

A commonly observed failure mode of Neural Radiance Field (NeRF) is fitting incorrect geometries when given an insufficient number of input views. One potential reason is that standard volumetric rendering does not enforce the constraint that most of a scene's geometry consist of empty space and opa…

Cited by 1030PDFcodeScholar
2022

Differentiable Raycasting for Self-Supervised Occupancy Forecasting

ECCV 2022poster

"Motion planning for safe autonomous driving requires learning how the environment around an ego-vehicle evolves with time. Ego-centric perception of driveable regions in a scene not only changes with the motion of actors in the environment, but also with the movement of the ego-vehicle itself. Self…

2022

Forecasting From LiDAR via Future Object Detection

CVPR 2022poster

Object detection and forecasting are fundamental components of embodied perception. These two problems, however, are largely studied in isolation by the community. In this paper, we propose an end-to-end approach for motion forecasting based on raw sensor measurement as opposed to ground truth track…

Cited by 40PDFcodeScholar
2022

HODOR: High-Level Object Descriptors for Object Re-Segmentation in Video Learned From Static Images

CVPR 2022oral

Existing state-of-the-art methods for Video Object Segmentation (VOS) learn low-level pixel-to-pixel correspondences between frames to propagate object masks across video. This requires a large amount of densely annotated video data, which is costly to annotate, and largely redundant since frames wi…

Cited by 30PDFcodeScholar
2022

Learning to Discover and Detect Objects

NeurIPS 2022accept

We tackle the problem of novel class discovery and localization (NCDL). In this setting, we assume a source dataset with supervision for only some object classes. Instances of other classes need to be discovered, classified, and localized automatically based on visual similarity without any human su…

2022

Mega-NERF: Scalable Construction of Large-Scale NeRFs for Virtual Fly-Throughs

CVPR 2022poster

We use neural radiance fields (NeRFs) to build interactive 3D environments from large-scale visual captures spanning buildings or even multiple city blocks collected primarily from drones. In contrast to single object scenes (on which NeRFs are traditionally evaluated), our scale poses multiple chal…

Cited by 414PDFcodeScholar
2022

Multimodal Object Detection via Probabilistic Ensembling

ECCV 2022poster

"Object detection with multimodal inputs can improve many safety-critical systems such as autonomous vehicles (AVs). Motivated by AVs that operate in both day and night, we study multimodal object detection with RGB and thermal cameras, since the latter provides much stronger object signatures under…

2022

RelPose: Predicting Probabilistic Relative Rotation for Single Objects in the Wild

ECCV 2022poster

"We describe a data-driven method for inferring the camera viewpoints given multiple images of an arbitrary object. This task is a core component of classic geometric pipelines such as SfM and SLAM, and also serves as a vital pre-processing requirement for contemporary neural approaches (e.g. NeRF)…

Cited by 94SourcePDFScholar
2021

Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

NeurIPS 2021poster

We introduce Argoverse 2 (AV2) — a collection of three datasets for perception and forecasting research in the self-driving domain. The annotated Sensor Dataset contains 1,000 sequences of multimodal data, encompassing high-resolution imagery from seven ring cameras, and two stereo cameras in additi…

Cited by 722SourcecodeScholar
2021

Background Splitting: Finding Rare Classes in a Sea of Background

CVPR 2021poster

We focus on the problem of training deep image classification models for a small number of extremely rare categories. In this common, real-world scenario, almost all images belong to the background category in the dataset. We find that state-of-the-art approaches for training on imbalanced datasets…

Cited by 8PDFcodeScholar
2021

Do Image Classifiers Generalize Across Time?

ICCV 2021poster

Vision models notoriously flicker when applied to videos: they correctly recognize objects in some frames, but fail on perceptually similar, nearby frames. In this work, we systematically analyze the robustness of image classifiers to such temporal perturbations in videos. To do so, we construct two…

Cited by 93PDFcodeScholar
2021

FOVEA: Foveated Image Magnification for Autonomous Navigation

ICCV 2021poster

Efficient processing of high-resolution video streams is safety-critical for many robotics applications such as autonomous driving. Image downsampling is a commonly adopted technique to ensure the latency constraint is met. However, this naive approach greatly restricts an object detector's capabili…

Cited by 40PDFcodeScholar
2021

LASR: Learning Articulated Shape Reconstruction From a Monocular Video

CVPR 2021poster

Remarkable progress has been made in 3D reconstruction of rigid structures from a video or a collection of images. However, it is still challenging to reconstruct nonrigid structures from RGB inputs, due to the under-constrained nature of this problem. While template-based approaches, such as parame…

Cited by 129PDFcodeScholar
2021

Learning Rare Category Classifiers on a Tight Labeling Budget

ICCV 2021poster

Many real-world ML deployments face the challenge of training a rare category model with a small labeling bud- get. In these settings, there is often access to large amounts of unlabeled data, therefore it is attractive to consider semi-supervised or active learning approaches to reduce human labeli…

Cited by 16PDFScholar
2021

Low-Shot Validation: Active Importance Sampling for Estimating Classifier Performance on Rare Categories

ICCV 2021poster

For machine learning models trained with limited labeled training data, validation stands to become the main bottleneck to reducing overall annotation costs. We propose a statistical validation algorithm that accurately estimates the F-score of binary classifiers for rare categories, where finding r…

Cited by 9PDFScholar
2021

NeRS: Neural Reflectance Surfaces for Sparse-view 3D Reconstruction in the Wild

NeurIPS 2021poster

Recent history has seen a tremendous growth of work exploring implicit representations of geometry and radiance, popularized through Neural Radiance Fields (NeRF). Such works are fundamentally based on a (implicit) {\em volumetric} representation of occupancy, allowing them to model diverse scene s…

2021

Safe Local Motion Planning With Self-Supervised Freespace Forecasting

CVPR 2021poster

Safe local motion planning for autonomous driving in dynamic environments requires forecasting how the scene evolves. Practical autonomy stacks adopt a semantic object-centric representation of a dynamic scene and build object detection, tracking, and prediction modules to solve forecasting. However…

Cited by 94PDFcodeScholar
2021

The CLEAR Benchmark: Continual LEArning on Real-World Imagery

NeurIPS 2021poster

Continual learning (CL) is widely regarded as crucial challenge for lifelong AI. However, existing CL benchmarks, e.g. Permuted-MNIST and Split-CIFAR, make use of artificial temporal variation and do not align with or generalize to the real- world. In this paper, we introduce CLEAR, the first contin…

Cited by 112SourcecodeScholar
2021

ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape Reconstruction

NeurIPS 2021spotlight

We introduce ViSER, a method for recovering articulated 3D shapes and dense3D trajectories from monocular videos. Previous work on high-quality reconstruction of dynamic 3D shapes typically relies on multiple camera views, strong category-specific priors, or 2D keypoint supervision. We show that no…

2020

4D Visualization of Dynamic Events From Unconstrained Multi-View Videos

CVPR 2020poster

We present a data-driven approach for 4D space-time visualization of dynamic events from videos captured by hand-held multiple cameras. Key to our approach is the use of self-supervised neural networks specific to the scene to compose static and dynamic aspects of an event. Though captured from disc…

Cited by 83PDFScholar
2020

Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints

ICLR 2020poster

In most practical settings and theoretical analyses, one assumes that a model can be trained until convergence. However, the growing complexity of machine learning datasets and models may violate such assumptions. Indeed, current approaches for hyper-parameter tuning and neural architecture search t…

Cited by 62SourceScholar
2020

Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild

ECCV 2020poster

We present a method that infers spatial arrangements and shapes of humans and objects in a globally consistent 3D scene, all from a single image in-the-wild captured in an uncontrolled environment. Notably, our method runs on datasets without any scene- or object-level 3D supervision. Our key insigh…

2020

TAO: A Large-Scale Benchmark for Tracking Any Object

ECCV 2020poster

For many years, multi-object tracking benchmarks have focused on a handful of categories. Motivated primarily by surveillance and self-driving applications, these datasets provide tracks for people, vehicles, and animals, ignoring the vast majority of objects in the world. By contrast, in the relate…

Cited by 214SourcePDFScholar
2020

What You See is What You Get: Exploiting Visibility for 3D Object Detection

CVPR 2020oral

Recent advances in 3D sensing have created unique challenges for computer vision. One fundamental challenge is finding a good representation for 3D sensor data. Most popular representations (such as PointNet) are proposed in the context of processing truly 3D data (e.g. points sampled from mesh mode…

Cited by 151PDFcodeScholar
2019

Argoverse: 3D Tracking and Forecasting With Rich Maps

CVPR 2019oral

We present Argoverse, a dataset designed to support autonomous vehicle perception tasks including 3D tracking and motion forecasting. Argoverse includes sensor data collected by a fleet of autonomous vehicles in Pittsburgh and Miami as well as 3D tracking annotations, 300k extracted interesting vehi…

Cited by 1736PDFcodeScholar
2019

DistInit: Learning Video Representations Without a Single Labeled Video

ICCV 2019poster

Video recognition models have progressed significantly over the past few years, evolving from shallow classifiers trained on hand-crafted features to deep spatiotemporal networks. However, labeled video data required to train such models has not been able to keep up with the ever increasing depth an…

Cited by 75PDFScholar
2019

Online Model Distillation for Efficient Video Inference

ICCV 2019poster

High-quality computer vision models typically address the problem of understanding the general distribution of real-world images. However, most cameras observe only a very small fraction of this distribution. This offers the possibility of achieving more efficient inference by specializing compact,…

Cited by 131PDFcodeScholar
2019

Weakly-Supervised Action Localization With Background Modeling

ICCV 2019poster

We describe a latent approach that learns to detect actions in long sequences given training videos with only whole-video class labels. Our approach makes use of two innovations to attention-modeling in weakly-supervised learning. First, and most notably, our framework uses an attention model to ext…

Cited by 208PDFcodeScholar
2018

Active Testing: An Efficient and Robust Framework for Estimating Accuracy

ICML 2018oral

Much recent work on large-scale visual recogni- tion aims to scale up learning to massive, noisily- annotated datasets. We address the problem of scaling-up the evaluation of such models to large- scale datasets with noisy labels. Current protocols for doing so require a human user to either vet (re…

Cited by 13SourcePDFScholar
2018

Few-Shot Human Motion Prediction via Meta-Learning

ECCV 2018poster

Human motion prediction, forecasting human motion in a few milliseconds conditioning on a historical 3D skeleton sequence, is a long-standing problem in computer vision and robotic vision. Existing forecasting algorithms rely on extensive annotated motion capture data and are brittle to novel action…

Cited by 155SourcePDFScholar
2017

ActionVLAD: Learning Spatio-Temporal Aggregation for Action Classification

CVPR 2017poster

In this work, we introduce a new video representation for action classification that aggregates local convolutional features across the entire spatio-temporal extent of the video. We do so by integrating state-of-the-art two-stream networks with learnable spatio-temporal feature aggregation. The res…

Cited by 607PDFScholar
2017

Expecting the Unexpected: Training Detectors for Unusual Pedestrians With Adversarial Imposters

CVPR 2017poster

As autonomous vehicles become an every-day reality, high-accuracy pedestrian detection is of paramount practical importance. Pedestrian detection is a highly researched topic with mature methods, but most datasets (for both training and evaluation) focus on common scenes of people engaged in typic…

Cited by 65PDFcodeScholar
2017

Finding Tiny Faces

CVPR 2017poster

Though tremendous strides have been made in object recognition, one of the remaining open challenges is detecting small objects. We explore three aspects of the problem in the context of finding small faces: the role of scale invariance, image resolution, and contextual reasoning. While most recogni…

Cited by 1025PDFScholar
2017

Need for Speed: A Benchmark for Higher Frame Rate Object Tracking

ICCV 2017poster

In this paper, we propose the first higher frame rate video dataset (called Need for Speed - NfS) and benchmark for visual object tracking. The dataset consists of 100 videos (380K frames) captured with now commonly available higher frame rate (240 FPS) cameras from real world scenarios. All frames…

Cited by 570PDFScholar
2017

Tracking as Online Decision-Making: Learning a Policy From Streaming Videos With Reinforcement Learning

ICCV 2017poster

We formulate tracking as an online decision-making process, where a tracking agent must follow an object despite ambiguous image frames and a limited computational budget. Crucially, the agent must decide where to look in the upcoming frames, when to reinitialize because it believes the target has b…

Cited by 140PDFScholar
2015

Depth-Based Hand Pose Estimation: Data, Methods, and Challenges

ICCV 2015poster

Hand pose estimation has matured rapidly in recent years. The introduction of commodity depth sensors and a multitude of practical applications have spurred new advances. We provide an extensive analysis of the state-of-the-art, focusing on hand pose estimation from a single depth frame. To do so, w…

Cited by 200PDFScholar
2015

Look and Think Twice: Capturing Top-Down Visual Attention With Feedback Convolutional Neural Networks

ICCV 2015poster

While feedforward deep convolutional neural networks (CNNs) have been a great success in computer vision, it is important to remember that the human visual contex contains generally more feedback connections than foward connections. In this paper, we will briefly introduce the background of feedback…

Cited by 530PDFcodeScholar