← Search

Sanja Fidler

163 accepted papers

2026

ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World Simulation

ICLR 2026poster

Recent advances in large generative models have significantly advanced image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, where edited objects must remain coherent. This capability is especially vital for world simulation related tasks. In thi…

Cited by 0SourcecodeScholar
2026

DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer

CVPR 2026

Simulation is essential to the development and evaluation of autonomous robots such as self-driving vehicles. Neural reconstruction is emerging as a promising solution as it enables simulating a wide variety of scenarios from real-world data alone in an automated and scalable way. However, while met

Cited by 0SourcecodeScholar
2026

Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation

ICLR 2026poster

The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the availability of captured real-world multi-view data, which is not…

Cited by 0SourcecodeScholar
2026

Motion Attribution for Video Generation

ICML 2026oral

Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework that scales to modern, large, high-quality video datasets and m…

Cited by 0SourceScholar
2026

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

ICLR 2026oral

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720×1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video alignment at a remarkably fast speed, deployable on RTX 5090 GPU. Tw…

Cited by 0SourcecodeScholar
2025

Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?

CVPR 2025poster

Improving semantic grounding in Vision-Language Models (VLMs) often involves collecting domain-specific training data, refining the network architectures, or modifying the training recipes. In this work, we venture into an orthogonal direction and explore self-correction in VLMs focusing on semantic…

Cited by 0SourcePDFScholar
2025

Controllable Weather Synthesis and Removal with Video Diffusion Models

ICCV 2025poster

Generating realistic and controllable weather effects in videos is valuable for many applications. Physics-based weather simulation requires precise reconstructions that are hard to scale to in-the-wild videos, while current video editing often lacks realism and control.In this work, we introduce We…

Cited by 0SourcePDFScholar
2025

DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models

CVPR 2025award

Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across representations. In this work, we introduce Difix3D+, a nov…

2025

Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion Models

CVPR 2025poster

Understanding and modeling lighting effects are fundamental tasks in computer vision and graphics. Classic physically-based rendering (PBR) accurately simulates the light transport, but relies on precise scene representations--explicit 3D geometry, high-quality material properties, and lighting cond…

Cited by 3SourcePDFScholar
2025

Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos

NeurIPS 2025poster

Recent advancements in static feed-forward scene reconstruction have demonstrated significant progress in high-quality novel view synthesis. However, these models often struggle with generalizability across diverse environments and fail to effectively handle dynamic content. We present BTimer (short…

Cited by 0SourceScholar
2025

GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control

CVPR 2025highlight

We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage little 3D information, leading to inconsistencies, such as objects popping in and out of existence. Camera control, if i…

2025

InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

ICCV 2025poster

We present InfiniCube, a scalable and controllable method to generate unbounded and dynamic 3D driving scenes with high fidelity.Previous methods for scene generation are constrained either by their applicability to indoor scenes or by their lack of controllability.In contrast, we take advantage of…

Cited by 0SourcePDFScholar
2025

LuxDiT: Lighting Estimation with Video Diffusion Transformer

NeurIPS 2025poster

Estimating scene lighting from a single image or video remains a longstanding challenge in computer vision and graphics. Learning-based approaches are constrained by the scarcity of ground-truth HDR environment maps, which are expensive to capture and limited in diversity. While recent generative mo…

Cited by 0SourceScholar
2025

OmniRe: Omni Urban Scene Reconstruction

ICLR 2025spotlight

We introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus on vehicles, hindering a holistic framework for all dynamic foregrounds demanded…

2025

PartField: Learning 3D Feature Fields for Part Segmentation and Beyond

ICCV 2025poster

We propose PartField, a feedforward approach for learning part-based 3D features, which captures the general concept of parts and their hierarchy without relying on predefined templates or text-based names, and can be applied to open-world 3D shapes across various modalities. PartField requires only…

Cited by 0SourcePDFScholar
2025

Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions

EMNLP 2025

Recent research in vision-language models (VLMs) has centered around the possibility of equipping them with implicit long-form chain-of-thought reasoning—akin to the success observed in language models—via distillation and reinforcement learning. But what about the non-reasoning models already train

Cited by 0SourcePDFScholar
2025

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

NeurIPS 2025spotlight

We address the challenge of relighting a single image or video, a task that demands precise scene intrinsic understanding and high-quality light transport synthesis. Existing end-to-end relighting models are often limited by the scarcity of paired multi-illumination data, restricting their ability t…

Cited by 0SourceScholar
2024

3DiffTection: 3D Object Detection with Geometry-Aware Diffusion Features

CVPR 2024poster

3DiffTection introduces a novel method for 3D object detection from single images utilizing a 3D-aware diffusion model for feature extraction. Addressing the resource-intensive nature of annotating large-scale 3D image data our approach leverages pretrained diffusion models traditionally used for 2D…

Cited by 13SourcePDFScholar
2024

Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models

CVPR 2024highlight

Text-guided diffusion models have revolutionized image and video generation and have also been successfully used for optimization-based 3D object synthesis. Here we instead focus on the underexplored text-to-4D setting and synthesize dynamic animated 3D objects using score distillation methods with…

Cited by 110SourcePDFScholar
2024

Align Your Steps: Optimizing Sampling Schedules in Diffusion Models

ICML 2024poster

Diffusion models (DMs) have established themselves as the state-of-the-art generative modeling approach in the visual domain and beyond. A crucial drawback of DMs is their slow sampling speed, relying on many sequential function evaluations through large neural networks. Sampling from DMs can be see…

Cited by 21SourcePDFScholar
2024

DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features

NeurIPS 2024poster

We propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving scenes. Our method is a generalizable feedforward model that predicts a rich neural scene representation from sparse, sing…

2024

EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models

ICLR 2024poster

Diffusion models have recently received increasing research attention for their remarkable transfer abilities in semantic segmentation tasks. However, generating fine-grained segmentation masks with diffusion models often requires additional training on annotated datasets, leaving it unclear to what…

2024

EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision

ICLR 2024poster

We present EmerNeRF, a simple yet powerful approach for learning spatial-temporal representations of dynamic driving scenes. Grounded in neural fields, EmerNeRF simultaneously captures scene geometry, appearance, motion, and semantics via self-bootstrapping. EmerNeRF hinges upon two core components:…

2024

L4GM: Large 4D Gaussian Reconstruction Model

NeurIPS 2024poster

We present L4GM, the first 4D Large Reconstruction Model that produces animated objects from a single-view video input -- in a single feed-forward pass that takes only a second. Key to our success is a novel dataset of multiview videos containing curated, rendered animated objects from Objaverse. Th…

Cited by 38SourcePDFScholar
2024

LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis

ECCV 2024poster

"Recent text-to-3D generation approaches produce impressive 3D results but require time-consuming optimization that can take up to an hour per prompt. Amortized methods like ATT3D optimize multiple prompts simultaneously to improve efficiency, enabling fast text-to-3D synthesis. However, they cannot…

2024

Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata

CVPR 2024highlight

We aim to generate fine-grained 3D geometry from large-scale sparse LiDAR scans abundantly captured by autonomous vehicles (AV). Contrary to prior work on AV scene completion we aim to extrapolate fine geometry from unlabeled and beyond spatial limits of LiDAR scans taking a step towards generating…

Cited by 0SourcePDFScholar
2024

Photorealistic Object Insertion with Diffusion-Guided Inverse Rendering

ECCV 2024poster

"The correct insertion of virtual objects in images of real-world scenes requires a deep understanding of the scene’s lighting, geometry and materials, as well as the image formation process. While recent large-scale diffusion models have shown strong generative and inpainting capabilities, we find…

Cited by 6SourcePDFScholar
2024

Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models

EMNLP 2024main

Despite recent advances demonstrating vision- language models’ (VLMs) abilities to describe complex relationships among objects in images using natural language, their capability to quantitatively reason about object sizes and distances remains underexplored. In this work, we introduce a manually an…

Cited by 7SourcePDFScholar
2024

SCube: Instant Large-Scale Scene Reconstruction using VoxSplats

NeurIPS 2024poster

We present SCube, a novel method for reconstructing large-scale 3D scenes (geometry, appearance, and semantics) from a sparse set of posed images. Our method encodes reconstructed scenes using a novel representation VoxSplat, which is a set of 3D Gaussians supported on a high-resolution sparse-voxel…

Cited by 10SourcePDFScholar
2024

Transferring Labels to Solve Annotation Mismatches Across Object Detection Datasets

ICLR 2024poster

In object detection, varying annotation protocols across datasets can result in annotation mismatches, leading to inconsistent class labels and bounding regions. Addressing these mismatches typically involves manually identifying common trends and fixing the corresponding bounding boxes and class la…

Cited by 1SourcePDFScholar
2024

WildFusion: Learning 3D-Aware Latent Diffusion Models in View Space

ICLR 2024poster

Modern learning-based approaches to 3D-aware image synthesis achieve high photorealism and 3D-consistent viewpoint changes for the generated images. Existing approaches represent instances in a shared canonical space. However, for in-the-wild datasets a shared canonical system can be difficult to de…

Cited by 6SourcePDFScholar
2024

XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies

CVPR 2024highlight

We present XCube a novel generative model for high-resolution sparse 3D voxel grids with arbitrary attributes. Our model can generate millions of voxels with a finest effective resolution of up to 1024^3 in a feed-forward fashion without time-consuming test-time optimization. To achieve this we empl…

2023

ATT3D: Amortized Text-to-3D Object Synthesis

ICCV 2023poster

Text-to-3D modelling has seen exciting progress by combining generative text-to-image models with image-to-3D methods like Neural Radiance Fields. DreamFusion recently achieved high-quality results but requires a lengthy, per-prompt optimization to create 3D objects. To address this, we amortize opt…

Cited by 82PDFScholar
2023

Align Your Latents: High-Resolution Video Synthesis With Latent Diffusion Models

CVPR 2023poster

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to high-resolution video generation, a particularly resource-intensive task. We fi…

2023

DreamTeacher: Pretraining Image Backbones with Deep Generative Models

ICCV 2023poster

In this work, we introduce a self-supervised feature representation learning framework DreamTeacher that utilizes generative networks for pre-training downstream image backbones. We propose to distill knowledge from a trained generative model into standard image backbones that have been well enginee…

Cited by 23PDFScholar
2023

End-to-end 3D Tracking with Decoupled Queries

ICCV 2023poster

In this work, we present an end-to-end framework for camera-based 3D multi-object tracking, called DQTrack. To avoid heuristic design in detection-based trackers, recent query-based approaches deal with identity-agnostic detection and identity-aware tracking in a single embedding. However, it brings…

Cited by 35PDFScholar
2023

Learning Human Dynamics in Autonomous Driving Scenarios

ICCV 2023poster

Simulation has emerged as an indispensable tool for scaling and accelerating the development of self-driving systems. A critical aspect of this is simulating realistic and diverse human behavior and intent. In this work, we propose a holistic framework for learning physically plausible human dynamic…

Cited by 23PDFScholar
2023

Magic3D: High-Resolution Text-to-3D Content Creation

CVPR 2023highlight

Recently, DreamFusion demonstrated the utility of a pretrained text-to-image diffusion model to optimize Neural Radiance Fields (NeRF), achieving remarkable text-to-3D synthesis results. However, the method has two inherent limitations: 1) optimization of the NeRF representation is extremely slow, 2…

Cited by 1196SourcePDFScholar
2023

Neural Fields Meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes

CVPR 2023poster

Reconstruction and intrinsic decomposition of scenes from captured imagery would enable many applications such as relighting and virtual object insertion. Recent NeRF based methods achieve impressive fidelity of 3D reconstruction, but bake the lighting and shadows into the radiance field, while mesh…

Cited by 89SourcePDFScholar
2023

Neural Kernel Surface Reconstruction

CVPR 2023highlight

We present a novel method for reconstructing a 3D implicit surface from a large-scale, sparse, and noisy point cloud. Our approach builds upon the recently introduced Neural Kernel Fields (NKF) representation. It enjoys similar generalization capabilities to NKF, while simultaneously addressing its…

Cited by 87SourcePDFScholar
2023

Neural LiDAR Fields for Novel View Synthesis

ICCV 2023poster

We present Neural Fields for LiDAR (NFL), a method to optimise a neural field scene representation from LiDAR measurements, with the goal of synthesizing realistic LiDAR scans from novel viewpoints. NFL combines the rendering power of neural fields with a detailed, physically motivated model of the…

Cited by 61PDFScholar
2023

NeuralField-LDM: Scene Generation With Hierarchical Latent Diffusion Models

CVPR 2023poster

Automatically generating high-quality real world 3D scenes is of enormous interest for applications such as virtual reality and robotics simulation. Towards this goal, we introduce NeuralField-LDM, a generative model capable of synthesizing complex 3D environments. We leverage Latent Diffusion Model…

2023

TexFusion: Synthesizing 3D Textures with Text-Guided Image Diffusion Models

ICCV 2023oral

We present TexFusion(Texture Diffusion), a new method to synthesize textures for given 3D geometries, using only large-scale text-guided image diffusion models. In contrast to recent works that leverage 2D text-to-image diffusion models to distill 3D objects using a slow and fragile optimization pro…

Cited by 99PDFcodeScholar
2023

Towards Viewpoint Robustness in Bird's Eye View Segmentation

ICCV 2023poster

Autonomous vehicles (AV) require that neural networks used for perception be robust to different viewpoints if they are to be deployed across many types of vehicles without the repeated cost of data collection and labeling for each. AV companies typically focus on collecting data from diverse scenar…

Cited by 15PDFScholar
2023

Trace and Pace: Controllable Pedestrian Animation via Guided Trajectory Diffusion

CVPR 2023poster

We introduce a method for generating realistic pedestrian trajectories and full-body animations that can be controlled to meet user-defined goals. We draw on recent advances in guided diffusion modeling to achieve test-time controllability of trajectories, which is normally only associated with rule…

Cited by 118SourcePDFScholar
2023

VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene Completion

CVPR 2023highlight

Humans can easily imagine the complete 3D geometry of occluded objects and scenes. This appealing ability is vital for recognition and understanding. To enable such capability in AI systems, we propose VoxFormer, a Transformer-based semantic scene completion framework that can output complete 3D vol…

2022

BigDatasetGAN: Synthesizing ImageNet With Pixel-Wise Annotations

CVPR 2022poster

Annotating images with pixel-wise labels is a time-consuming and costly process. Recently, DatasetGAN showcased a promising alternative - to synthesize a large labeled dataset via a generative adversarial network (GAN) by exploiting a small set of manually labeled, GAN-generated images. Here, we sca…

Cited by 121PDFScholar
2022

EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations

NeurIPS 2022accept

We introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets. Specifically, we need…

2022

Extracting Triangular 3D Models, Materials, and Lighting From Images

CVPR 2022oral

We present an efficient method for joint optimization of topology, materials and lighting from multi-view image observations. Unlike recent multi-view reconstruction approaches, which typically produce entangled 3D representations encoded in neural networks, we output triangle meshes with spatially-…

Cited by 404PDFcodeScholar
2022

GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images

NeurIPS 2022accept

As several industries are moving towards modeling massive 3D virtual worlds, the need for content creation tools that can scale in terms of the quantity, quality, and diversity of 3D content is becoming evident. In our work, we aim to train performant 3D generative models that synthesize textured me…

2022

Generating Useful Accident-Prone Driving Scenarios via a Learned Traffic Prior

CVPR 2022poster

Evaluating and improving planning for autonomous vehicles requires scalable generation of long-tail traffic scenarios. To be useful, these scenarios must be realistic and challenging, but not impossible to drive through safely. In this work, we introduce STRIVE, a method to automatically generate ch…

Cited by 157PDFScholar
2022

How Much More Data Do I Need? Estimating Requirements for Downstream Tasks

CVPR 2022poster

Given a small training data set and a learning algorithm, how much more data is necessary to reach a target validation or test performance? This question is of critical importance in applications such as autonomous driving or medical imaging where collecting data is expensive and time-consuming. Ove…

Cited by 32PDFScholar
2022

LION: Latent Point Diffusion Models for 3D Shape Generation

NeurIPS 2022accept

Denoising diffusion models (DDMs) have shown promising results in 3D point cloud synthesis. To advance 3D DDMs and make them useful for digital artists, we require (i) high generation quality, (ii) flexibility for manipulation and applications such as conditional synthesis and shape interpolation, a…

2022

Low-Budget Active Learning via Wasserstein Distance: An Integer Programming Approach

ICLR 2022poster

Active learning is the process of training a model with limited labeled data by selecting a core subset of an unlabeled data pool to label. The large scale of data sets used in deep learning forces most sample selection strategies to employ efficient heuristics. This paper introduces an integer opti…

Cited by 45SourcePDFScholar
2022

MvDeCor: Multi-View Dense Correspondence Learning for Fine-Grained 3D Segmentation

ECCV 2022poster

"We propose to utilize self-supervised techniques in the 2D domain for fine-grained 3D shape segmentation tasks. This is inspired by the observation that view-based surface representations are more effective at modeling high-resolution surface details and texture than their 3D counterparts based on…

Cited by 13SourcePDFScholar
2022

Neural Fields As Learnable Kernels for 3D Reconstruction

CVPR 2022poster

We present Neural Kernel Fields: a novel method for reconstructing implicit 3D shapes based on a learned kernel ridge regression. Our technique achieves state-of-the-art results when reconstructing 3D objects and large scenes from sparse oriented points, and can reconstruct shape categories outside…

Cited by 82PDFScholar
2022

Neural Light Field Estimation for Street Scenes with Differentiable Virtual Object Insertion

ECCV 2022poster

"We consider the challenging problem of outdoor lighting estimation for the goal of photorealistic virtual object insertion into photographs. Existing works on outdoor lighting estimation typically simplify the scene lighting into an environment map which cannot capture the spatially-varying lightin…

Cited by 41SourcePDFScholar
2022

Optimizing Data Collection for Machine Learning

NeurIPS 2022accept

Modern deep learning systems require huge data sets to achieve impressive performance, but there is little guidance on how much or what kind of data to collect. Over-collecting data incurs unnecessary present costs, while under-collecting may incur future costs and delay workflows. We propose a new…

Cited by 38SourcePDFScholar
2022

Polymorphic-GAN: Generating Aligned Samples Across Multiple Domains With Learned Morph Maps

CVPR 2022oral

Modern image generative models show remarkable sample quality when trained on a single domain or class of objects. In this work, we introduce a generative adversarial network that can simultaneously generate aligned image samples from multiple related domains. We leverage the fact that a variety of…

Cited by 9PDFScholar
2021

3DStyleNet: Creating 3D Shapes With Geometric and Texture Style Variations

ICCV 2021poster

We propose a method to create plausible geometric and texture style variations of 3D objects in the quest to democratize 3D content creation. Given a pair of textured source and target objects, our method predicts a part-aware affine transformation field that naturally warps the source shape to imit…

Cited by 75PDFScholar
2021

ATISS: Autoregressive Transformers for Indoor Scene Synthesis

NeurIPS 2021poster

The ability to synthesize realistic and diverse indoor furniture layouts automatically or based on partial input, unlocks many applications, from better interactive 3D tools to data synthesis for training and simulation. In this paper, we present ATISS, a novel autoregressive transformer architectur…

2021

DIB-R++: Learning to Predict Lighting and Material with a Hybrid Differentiable Renderer

NeurIPS 2021poster

We consider the challenging problem of predicting intrinsic object properties from a single image by exploiting differentiable renderers. Many previous learning-based approaches for inverse graphics adopt rasterization-based renderers and assume naive lighting and material models, which often fail t…

Cited by 68SourcePDFScholar
2021

DatasetGAN: Efficient Labeled Data Factory With Minimal Human Effort

CVPR 2021poster

We introduce DatasetGAN: an automatic procedure to generate massive datasets of high-quality semantically segmented images requiring minimal human effort. Current deep networks are extremely data-hungry, benefiting from training on large-scale datasets, which are time-consuming to annotate. Our meth…

Cited by 393PDFcodeScholar
2021

Deep Marching Tetrahedra: a Hybrid Representation for High-Resolution 3D Shape Synthesis

NeurIPS 2021poster

We introduce DMTet, a deep 3D conditional generative model that can synthesize high-resolution 3D shapes using simple user guides such as coarse voxels. It marries the merits of implicit and explicit 3D representations by leveraging a novel hybrid 3D representation. Compared to the current implicit…

2021

Don’t Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence

NeurIPS 2021poster

Although machine learning models trained on massive data have led to breakthroughs in several areas, their deployment in privacy-sensitive domains remains limited due to restricted access to data. Generative models trained with privacy constraints on private data can sidestep this challenge, providi…

Cited by 83SourcePDFScholar
2021

DriveGAN: Towards a Controllable High-Quality Neural Simulation

CVPR 2021poster

Realistic simulators are critical for training and verifying robotics systems. While most of the contemporary simulators are hand-crafted, a scaleable way to build simulators is to use machine learning to learn how the environment behaves in response to an action, directly from data. In this work, w…

Cited by 119PDFScholar
2021

EditGAN: High-Precision Semantic Image Editing

NeurIPS 2021poster

Generative adversarial networks (GANs) have recently found applications in image editing. However, most GAN-based image editing methods often require large-scale datasets with semantic segmentation annotations for training, only provide high-level control, or merely interpolate between different ima…

Cited by 280SourcePDFScholar
2021

Emergent Road Rules In Multi-Agent Driving Environments

ICLR 2021poster

For autonomous vehicles to safely share the road with human drivers, autonomous vehicles must abide by specific "road rules" that human drivers have agreed to follow. "Road rules" include rules that drivers are required to follow by law – such as the requirement that vehicles stop at red lights – as…

2021

Image GANs meet Differentiable Rendering for Inverse Graphics and Interpretable 3D Neural Rendering

ICLR 2021oral

Differentiable rendering has paved the way to training neural networks to perform “inverse graphics” tasks such as predicting 3D geometry from monocular photographs. To train high performing models, most of the current approaches rely on multi-view imagery which are not readily available in practice…

Cited by 145SourcePDFScholar
2021

Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection

ICML 2021spotlight

Training on datasets with long-tailed distributions has been challenging for major recognition tasks such as classification and detection. To deal with this challenge, image resampling is typically introduced as a simple but effective approach. However, we observe that long-tailed detection differs…

2021

NP-DRAW: A Non-Parametric Structured Latent Variable Model for Image Generation

UAI 2021poster

In this paper, we present a non-parametric structured latent variable model for image generation, called NP-DRAW, which sequentially draws on a latent canvas in a part-by-part fashion and then decodes the image from the canvas. Our key contributions are as follows. 1) We propose a non-parametric pri…

2021

Neural Geometric Level of Detail: Real-Time Rendering With Implicit 3D Shapes

CVPR 2021poster

Neural signed distance functions (SDFs) are emerging as an effective representation for 3D shapes. State-of-the-art methods typically encode the SDF with a large, fixed-size neural network to approximate complex shapes with implicit surfaces. Rendering with these large networks is, however, computat…

Cited by 543PDFcodeScholar
2021

Neural Parts: Learning Expressive 3D Shape Abstractions With Invertible Neural Networks

CVPR 2021poster

Impressive progress in 3D shape extraction led to representations that can capture object geometries with high fidelity. In parallel, primitive-based methods seek to represent objects as semantically consistent part arrangements. However, due to the simplicity of existing primitive representations,…

Cited by 123PDFcodeScholar
2021

Personalized Federated Learning with First Order Model Optimization

ICLR 2021poster

While federated learning traditionally aims to train a single global model across decentralized local datasets, one model may not always be ideal for all participating clients. Here we propose an alternative, where each client only federates with other relevant clients to obtain a stronger model per…

2021

Physics-Based Human Motion Estimation and Synthesis From Videos

ICCV 2021poster

Human motion synthesis is an important problem for applications in graphics and gaming, and even in simulation environments for robotics. Existing methods require accurate motion capture data for training, which is costly to obtain. Instead, we propose a framework for training generative models of p…

Cited by 110PDFScholar
2021

Scalable Neural Data Server: A Data Recommender for Transfer Learning

NeurIPS 2021poster

Absence of large-scale labeled data in the practitioner's target domain can be a bottleneck to applying machine learning algorithms in practice. Transfer learning is a popular strategy for leveraging additional data to improve the downstream performance, but finding the most relevant data to transfe…

Cited by 8SourcePDFScholar
2021

Semantic Segmentation With Generative Models: Semi-Supervised Learning and Strong Out-of-Domain Generalization

CVPR 2021poster

Training deep networks with limited labeled data while achieving a strong generalization ability is key in the quest to reduce human annotation efforts. This is the goal of semi-supervised learning, which exploits more widely available unlabeled data to complement small labeled data sets. In this pa…

Cited by 241PDFcodeScholar
2021

Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets

CVPR 2021poster

Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation strategies for collecting multi-class classification labels for a l…

Cited by 37PDFScholar
2021

Towards Optimal Strategies for Training Self-Driving Perception Models in Simulation

NeurIPS 2021poster

Autonomous driving relies on a huge volume of real-world data to be labeled to high precision. Alternative solutions seek to exploit driving simulators that can generate large amounts of labeled data with a plethora of content variations. However, the domain gap between the synthetic and real data…

Cited by 23SourcePDFScholar
2021

Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration

ICLR 2021spotlight

In this paper, we introduce Watch-And-Help (WAH), a challenge for testing social intelligence in agents. In WAH, an AI agent needs to help a human-like agent perform a complex household task efficiently. To succeed, the AI agent needs to i) understand the underlying goal of the task by watching a si…

2021

gradSim: Differentiable simulation for system identification and visuomotor control

ICLR 2021poster

In this paper, we tackle the problem of estimating object physical properties such as mass, friction, and elasticity directly from video sequences. Such a system identification problem is fundamentally ill-posed due to the loss of information during image formation. Current best solutions to the pro…

Cited by 40SourcePDFScholar
2020

Auto-Tuning Structured Light by Optical Stochastic Gradient Descent

CVPR 2020poster

We consider the problem of optimizing the performance of an active imaging system by automatically discovering the illuminations it should use, and the way to decode them. Our approach tackles two seemingly incompatible goals: (1) "tuning" the illuminations and decoding algorithm precisely to the de…

Cited by 31PDFcodeScholar
2020

Beyond Fixed Grid: Learning Geometric Image Representation with a Deformable Grid

ECCV 2020poster

In modern computer vision, images are typically represented as a fixed uniform grid with some stride and processed via a deep convolutional neural network. We argue that deforming the grid to better align with the high-frequency image content is a more effective strategy. We introduce mph{Deformable…

2020

Efficient and Information-Preserving Future Frame Prediction and Beyond

ICLR 2020poster

Applying resolution-preserving blocks is a common practice to maximize information preservation in video prediction, yet their high memory consumption greatly limits their application scenarios. We propose CrevNet, a Conditionally Reversible Network that uses reversible architectures to build a bije…

Cited by 142SourceScholar
2020

Expressive Telepresence via Modular Codec Avatars

ECCV 2020poster

VR telepresence consists of interacting with another human in a virtual space represented by an avatar. Today most avatars are cartoon-like, but soon the technology will allow video-realistic ones. This paper aims in this direction and presents Modular Codec Avatars (MCA), a method to generate hyper…

Cited by 40SourcePDFScholar
2020

Interactive Annotation of 3D Object Geometry using 2D Scribbles

ECCV 2020poster

Inferring detailed 3D geometry of the scene is crucial for robotics applications, simulation, and 3D content creation. However, such information is hard to obtain, and thus very few datasets support it. In this paper, we propose an interactive framework for annotating 3D object geometry from both po…

Cited by 17SourcePDFScholar
2020

Learning Deformable Tetrahedral Meshes for 3D Reconstruction

NeurIPS 2020poster

3D shape representations that accommodate learning-based 3D reconstruction are an open problem in machine learning and computer graphics. Previous work on neural 3D reconstruction demonstrated benefits, but also limitations, of point cloud, voxel, surface mesh, and implicit function representations.…

2020

Learning to Simulate Dynamic Environments With GameGAN

CVPR 2020poster

Simulation is a crucial component of any robotic system. In order to simulate correctly, we need to write complex rules of the environment: how dynamic agents behave, and how the actions of each of the agents affect the behavior of others. In this paper, we aim to learn a simulator by simply watchin…

Cited by 138PDFScholar
2020

Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D

ECCV 2020poster

Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D","The goal of perception for autonomous vehicles is to extract semantic representations from multiple sensors and fuse these representations into a single ``bird's-eye-view'' coordinate frame for consumption by…

2020

Meta-Sim2: Unsupervised Learning of Scene Structure for Synthetic Data Generation

ECCV 2020poster

Generation of synthetic data has allowed Machine Learning practitioners to bypass the need for costly collection and labeling of large datasets. Unfortunately the generation of such data often requires experts to carefully design sampling procedures that guarantee creation of realistic scenes. These…

Cited by 109SourcePDFScholar
2020

ScribbleBox: Interactive Annotation Framework for Video Object Segmentation

ECCV 2020poster

Manually labeling video datasets for segmentation tasks is extremely time consuming. We introduce ScribbleBox, an interactive framework for annotating object instances with masks in videos with a significant boost in efficiency. In particular, we split annotation into two steps: annotating objects w…

Cited by 23SourcePDFScholar
2019

Action Recognition From Single Timestamp Supervision in Untrimmed Videos

CVPR 2019poster

Recognising actions in videos relies on labelled supervision during training, typically the start and end times of each action instance. This supervision is not only subjective, but also expensive to acquire. Weak video-level supervision has been successfully exploited for recognition in untrimmed v…

Cited by 92PDFcodeScholar
2019

DMM-Net: Differentiable Mask-Matching Network for Video Object Segmentation

ICCV 2019poster

In this paper, we propose the differentiable mask-matching network (DMM-Net) for solving the video object segmentation problem where the initial object masks are provided. Relying on the Mask R-CNN backbone, we extract mask proposals per frame and formulate the matching between object templates and…

Cited by 98PDFcodeScholar
2019

EigenDamage: Structured Pruning in the Kronecker-Factored Eigenbasis

ICML 2019oral

Reducing the test time resource requirements of a neural network while preserving test accuracy is crucial for running inference on resource-constrained devices. To achieve this goal, we introduce a novel network reparameterization based on the Kronecker-factored eigenbasis (KFE), and then apply Hes…

2019

Learning to Predict 3D Objects with an Interpolation-based Differentiable Renderer

NeurIPS 2019poster

Many machine learning models operate on images, but ignore the fact that images are 2D projections formed by 3D geometry interacting with light, in a process called rendering. Enabling ML models to understand image formation might be key for generalization. However, due to an essential rasterization…

Cited by 453SourcePDFScholar
2019

Meta-Sim: Learning to Generate Synthetic Datasets

ICCV 2019oral

Training models to high-end performance requires availability of large labeled datasets, which are expensive to get. The goal of our work is to automatically synthesize labeled datasets that are relevant for a downstream task. We propose Meta-Sim, which learns a generative model of synthetic scenes,…

Cited by 316PDFScholar
2019

Neural Graph Evolution: Towards Efficient Automatic Robot Design

ICLR 2019poster

Despite the recent successes in robotic locomotion control, the design of robot relies heavily on human engineering. Automatic robot design has been a long studied subject, but the recent progress has been slowed due to the large combinatorial search space and the difficulty in evaluating the found…

2019

Neural Turtle Graphics for Modeling City Road Layouts

ICCV 2019oral

We propose Neural Turtle Graphics (NTG), a novel generative model for spatial graphs, and demonstrate its applications in modeling city road layouts. Specifically, we represent the road layout using a graph where nodes in the graph represent control points and edges in the graph represents road segm…

Cited by 104PDFScholar
2019

Object Instance Annotation With Deep Extreme Level Set Evolution

CVPR 2019poster

In this paper, we tackle the task of interactive object segmentation. We revive the old ideas on level set segmentation which framed object annotation as curve evolution. Carefully designed energy functions ensured that the curve was well aligned with image boundaries, and generally "well behaved".…

Cited by 98PDFcodeScholar
2019

Synthesizing Environment-Aware Activities via Activity Sketches

CVPR 2019poster

In order to learn to perform activities from demonstrations or descriptions, agents need to distill what the essence of the given activity is, and how it can be adapted to new environments. In this work, we address the problem: environment-aware program generation. Given a visual demonstration or a…

Cited by 45PDFScholar
2018

Efficient Interactive Annotation of Segmentation Datasets With Polygon-RNN++

CVPR 2018poster

Manually labeling datasets with object masks is extremely time consuming. In this work, we follow the idea of Polygon-RNN to produce polygonal annotations of objects interactively using humans-in-the-loop. We introduce several important improvements to the model: 1) we design a new CNN encoder archi…

Cited by 537SourcePDFScholar
2018

Learning to Act Properly: Predicting and Explaining Affordances From Images

CVPR 2018poster

We address the problem of affordance reasoning in diverse scenes that appear in the real world. Affordances relate the agent’s actions to their effects when taken on the surrounding objects. In our work, we take the egocentric view of the scene, and aim to reason about action-object affordances that…

Cited by 125SourcePDFScholar
2018

MovieGraphs: Towards Understanding Human-Centric Situations From Videos

CVPR 2018poster

There is growing interest in artificial intelligence to build socially intelligent robots. This requires machines to have the ability to "read" people's emotions, motivations, and other factors that affect behavior. Towards this goal, we introduce a novel dataset called MovieGraphs which provides de…

Cited by 179SourcePDFScholar
2018

NerveNet: Learning Structured Policy with Graph Neural Networks

ICLR 2018poster

We address the problem of learning structured policies for continuous control. In traditional reinforcement learning, policies of agents are learned by MLPs which take the concatenation of all observations from the environment as input for predicting actions. In this work, we propose NerveNet to exp…

Cited by 326SourcePDFScholar
2018

Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

ECCV 2018poster

First-person vision is gaining interest as it offers a unique viewpoint on people’s interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due to the lack of sufficiently large datasets. In this paper, we introduce EPIC-KI…

Cited by 1329SourcePDFScholar
2018

SurfConv: Bridging 3D and 2D Convolution for RGBD Images

CVPR 2018poster

The last few years have seen approaches trying to combine the increasing popularity of depth sensors and the success of the convolutional neural networks. Using depth as additional channel alongside the RGB input has the scale variance problem present in image convolution based approaches. On the ot…

2018

VirtualHome: Simulating Household Activities via Programs

CVPR 2018poster

In this paper, we are interested in modeling complex activities that occur in a typical household. We propose to use programs, i.e., sequences of atomic actions and interactions, as a high level representation of complex tasks. Programs are interesting because they provide a non-ambiguous representa…

Cited by 620SourcePDFScholar
2017

3D Graph Neural Networks for RGBD Semantic Segmentation

ICCV 2017oral

RGBD semantic segmentation requires joint reasoning about 2D appearance and 3D geometric information. In this paper we propose a 3D graph neural network (3DGNN) that builds a k-nearest neighbor graph on top of 3D point cloud. Each node in the graph corresponds to a set of points and is associated wi…

Cited by 605PDFcodeScholar
2017

Be Your Own Prada: Fashion Synthesis With Structural Coherence

ICCV 2017poster

We present a novel and effective approach for generating new clothing on a wearer through generative adversarial learning. Given an input image of a person and a sentence describing a different outfit, our model "redresses" the person as desired, while at the same time keeping the wearer and her/his…

Cited by 346PDFScholar
2017

Find your way by observing the sun and other semantic cues

ICRA 2017poster

In this paper we present a robust, efficient and affordable approach to self-localization which requires neither GPS nor knowledge about the appearance of the world. Towards this goal, we utilize freely available cartographic maps and derive a probabilistic model that exploits semantic cues in the f…

Cited by 60SourceScholar
2017

Situation Recognition With Graph Neural Networks

ICCV 2017poster

We address the problem of recognizing situations in images. Given an image, the task is to predict the most salient verb (action), and fill its semantic roles such as who is performing the action, what is the source and target of the action, etc. Different verbs have different roles (e.g. attacking…

Cited by 142PDFScholar
2017

TorontoCity: Seeing the World With a Million Eyes

ICCV 2017spotlight

In this paper we introduce the TorontoCity benchmark, which covers the full greater Toronto area (GTA) with 712.5km2 of land, 8439km of road and around 400, 000 buildings. Our benchmark provides different perspectives of the world captured from airplanes, drones and cars driving around the city. Man…

Cited by 217PDFScholar
2017

Towards Diverse and Natural Image Descriptions via a Conditional GAN

ICCV 2017oral

Despite the substantial progress in recent years, the problem of image captioning remains far from being satisfactorily tackled. Sentences produced by existing methods, e.g. those based on LSTM, are often overly rigid and lacking in variability. This issue is related to a learning principle widely u…

Cited by 804PDFcodeScholar
2016

HD Maps: Fine-Grained Road Segmentation by Parsing Ground and Aerial Images

CVPR 2016poster

In this paper we present an approach to enhance existing maps with fine grained segmentation categories such as parking spots and sidewalk, as well as the number and location of road lanes. Towards this goal, we propose an efficient approach that is able to estimate these fine grained categories by…

Cited by 181PDFScholar
2016

Instance-Level Segmentation for Autonomous Driving With Deep Densely Connected MRFs

CVPR 2016poster

Our aim is to provide a pixel-wise instance-level labeling of a monocular image in the context of autonomous driving. We build on recent work [Zhang et al., ICCV15] that trained a convolutional neural net to predict instance labeling in local image patches, extracted exhaustively in a stride from an…

Cited by 292PDFScholar
2016

Monocular 3D Object Detection for Autonomous Driving

CVPR 2016poster

The goal of this paper is to perform 3D object detection in single monocular images in the domain of autonomous driving. Our method first aims to generate a set of candidate class-specific object proposals, which are then run through a standard CNN pipeline to obtain high-quality object detections.…

Cited by 1263PDFScholar
2016

MovieQA: Understanding Stories in Movies Through Question-Answering

CVPR 2016spotlight

We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who" did "What" to "Whom", to "Why" and "How" certain events occur…

Cited by 875PDFScholar
2015

3D Object Proposals for Accurate Object Class Detection

NeurIPS 2015poster

The goal of this paper is to generate high-quality 3D object proposals in the context of autonomous driving. Our method exploits stereo imagery to place proposals in the form of 3D bounding boxes. We formulate the problem as minimizing an energy function encoding object size priors, ground plane a…

Cited by 1092SourcePDFScholar
2015

Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books

ICCV 2015oral

Books are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This paper aims to align books to their movie releases in order to provide rich…

Cited by 3512PDFScholar
2015

Monocular Object Instance Segmentation and Depth Ordering With CNNs

ICCV 2015poster

In this paper we tackle the problem of instance-level segmentation and depth ordering from a single monocular image. Towards this goal, we take advantage of convolutional neural nets and train them to directly predict instance-level segmentations where the instance ID encodes the depth ordering with…

Cited by 194PDFScholar
2015

Neuroaesthetics in Fashion: Modeling the Perception of Fashionability

CVPR 2015poster

In this paper, we analyze the fashion of clothing of a large social website. Our goal is to learn and predict how fashionable a person looks on a photograph and suggest subtle improvements the user could make to improve her/his appeal. We propose a Conditional Random Field model that jointly reasons…

Cited by 251SourcePDFScholar
2015

Predicting Deep Zero-Shot Convolutional Neural Networks Using Textual Descriptions

ICCV 2015poster

One of the main challenges in Zero-Shot Learning of visual categories is gathering semantic attributes to accompany images. Recent work has shown that learning from textual descriptions, such as Wikipedia articles, avoids the problem of having to explicitly define these attributes. We present a new…

Cited by 527PDFScholar
2015

Real-Time Coarse-to-Fine Topologically Preserving Segmentation

CVPR 2015poster

In this paper, we tackle the problem of unsupervised segmentation in the form of superpixels. Our main emphasis is on speed and accuracy. We build on [31] to define the problem as a boundary and topology preserving Markov random field. We propose a coarse to fine optimization technique that speeds u…

Cited by 177SourcePDFScholar
2015

Rent3D: Floor-Plan Priors for Monocular Layout Estimation

CVPR 2015poster

The goal of this paper is to enable a 3D "virtual-tour" of an apartment given a small set of monocular images of different rooms, as well as a 2D floor plan. We frame the problem as inference in a Markov Random Field which reasons about the layout of each room and its relative pose (3D rotation and…

Cited by 177SourcePDFScholar
2015

Skip-Thought Vectors

NeurIPS 2015poster

We describe an approach for unsupervised learning of a generic, distributed sentence encoder. Using the continuity of text from books, we train an encoder-decoder model that tries to reconstruct the surrounding sentences of an encoded passage. Sentences that share semantic and syntactic properties a…

2015

segDeepM: Exploiting Segmentation and Context in Deep Neural Networks for Object Detection

CVPR 2015poster

In this paper, we propose an approach that exploits object segmentation in order to improve the accuracy of object detection. We frame the problem as inference in a Markov Random Field, in which each detection hypothesis scores object appearance as well as contextual information using Convolutional…

Cited by 211SourcePDFScholar