← Search

Andreas Geiger

122 accepted papers

2026

3D-LATTE: Latent Space 3D Editing from Textual Instructions

CVPR 2026

Despite the recent success of multi-view diffusion models for text/image-based 3D asset generation, instruction-based editing of 3D assets lacks surprisingly far behind the quality of generation models. The main reason is that recent approaches using 2D priors suffer from view-inconsistent editing s

Cited by 0SourcecodeScholar
2026

EMPERROR: A Flexible Generative Perception Error Model for Probing Self-Driving Planners

ICRA 2026poster

To handle the complexities of real-world traffic, learning planners for self-driving from data is a promising direction. While recent approaches have shown great progress, they typically assume a setting in which the ground-truth world state is available as input. However, when deployed, planning ne…

2026

FrankenMotion: Part-level Human Motion Generation and Composition

CVPR 2026

Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptions due to the absence of fine-grained, part-level motion annotations. This limits their controllability over individual

Cited by 0SourcecodeScholar
2026

LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving

CVPR 2026

Simulators can generate virtually unlimited driving data, yet imitation learning policies in simulation still struggle to achieve robust closed-loop performance. Motivated by this gap, we empirically study how misalignment between privileged expert demonstrations and sensor-based student observation

Cited by 0SourcecodeScholar
2026

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

ICML 2026poster

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures or necessitate compressing geometry into latent spaces to leverage pre-trained latent diffusion models. In this work, we demonstrate that such architectural overhead is unnecessary. We introduce a mini…

Cited by 0SourceScholar
2026

PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban Scenes

CVPR 2026

Existing approaches to 3D semantic urban scene generation predominantly rely on voxel-based representations, which are bound by fixed resolution, challenging to edit, and memory-intensive in their dense form. In contrast, we advocate for a primitive-based paradigm where urban scenes are represented

Cited by 0SourcecodeScholar
2026

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

CVPR 2026

End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-gra

Cited by 0SourcecodeScholar
2025

CaRL: Learning Scalable Planning Policies with Simple Rewards

CoRL 2025poster

We investigate reinforcement learning (RL) for privileged planning in autonomous driving. State-of-the-art approaches for this task are rule-based, but these methods do not scale to the long tail. RL, on the other hand, is scalable and does not suffer from compounding errors like imitation learning…

Cited by 0SourcecodeScholar
2025

DepthSplat: Connecting Gaussian Splatting and Depth

CVPR 2025poster

Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present DepthSplat to connect Gaussian splatting and depth estimation and study their interactions. More specifically, we first contribute a robust multi-view depth model by leveraging pre-train…

2025

EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View Synthesis

CVPR 2025poster

Novel view synthesis of urban scenes is essential for autonomous driving-related applications. Existing NeRF and 3DGS-based methods show promising results in achieving photorealistic renderings but require slow, per-scene optimization. We introduce EVolSplat, an efficient 3D Gaussian Splatting model…

Cited by 0SourcePDFScholar
2025

Easi3R: Estimating Disentangled Motion from DUSt3R Without Training

ICCV 2025poster

Recent advances in DUSt3R have enabled robust estimation of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct supervision on large-scale 3D datasets.In contrast, the limited scale and diversity of available 4D datasets present a major…

2025

Emperror: A Flexible Generative Perception Error Model for Probing Self-Driving Planners

RA-L 2025

To handle the complexities of real-world traffic, learning planners for self-driving from data is a promising direction. While recent approaches have shown great progress, they typically assume a setting in which the ground-truth world state is available as input. However, when deployed, planning ne

Cited by 2SourceScholar
2025

GenFusion: Closing the Loop between Reconstruction and Generation via Videos

CVPR 2025poster

Recently, 3D reconstruction and generation have demonstrated impressive novel view synthesis results, achieving high fidelity and efficiency. However, a notable conditioning gap can be observed between these two fields, e.g. , scalable 3D scene reconstruction often requires densely captured views,…

Cited by 0SourcePDFScholar
2025

LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models

ICCV 2025poster

Vision foundation models (VFMs) such as DINOv2 and CLIP have achieved impressive results on various downstream tasks, but their limited feature resolution hampers performance in applications requiring pixel-level understanding. Feature upsampling offers a promising direction to address this challeng…

2025

MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction

ICCV 2025poster

We present MoGA, a novel method to reconstruct high-fidelity 3D Gaussian avatars from a single-view image. The main challenge lies in inferring unseen appearance and geometric details while ensuring 3D consistency and realism. Most previous methods rely on 2D diffusion models to synthesize unseen vi…

2025

Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

CVPR 2025poster

In this work, we introduce Prometheus, a 3D-aware latent diffusion model for text-to-3D generation at both object and scene levels in seconds. We formulate 3D scene generation as multi-view, feed-forward, pixel-aligned 3D Gaussian generation within the latent diffusion paradigm. To ensure generaliza…

Cited by 3SourcePDFScholar
2025

Pseudo-Simulation for Autonomous Driving

CoRL 2025poster

Existing evaluation paradigms for Autonomous Vehicles (AVs) face critical limitations. Real-world evaluation is often challenging due to safety concerns and a lack of reproducibility, whereas closed-loop simulation can face insufficient realism or high computational costs. Open-loop evaluation, whil…

Cited by 0SourcecodeScholar
2025

ReSim: Reliable World Simulation for Autonomous Driving

NeurIPS 2025spotlight

How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusively on real-world driving data composed mainly of safe expert trajectories, struggle to follow hazardous or non-expert behaviors, which are rare in such d…

Cited by 0SourceScholar
2025

UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation

CVPR 2025poster

Photorealistic 3D vehicle models with high controllability are essential for autonomous driving simulation and data augmentation. While handcrafted CAD models provide flexible controllability, free CAD libraries often lack the high-quality materials necessary for photorealistic rendering. Conversely…

Cited by 0SourcePDFScholar
2025

Volumetric Surfaces: Representing Fuzzy Geometries with Layered Meshes

CVPR 2025poster

High-quality view synthesis relies on volume rendering, splatting, or surface rendering. While surface rendering is typically the fastest, it struggles to accurately model fuzzy geometry like hair. In turn, alpha-blending techniques excel at representing fuzzy materials but require an unbounded numb…

2024

3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting

CVPR 2024poster

We introduce an approach that creates animatable human avatars from monocular videos using 3D Gaussian Splatting (3DGS). Existing methods based on neural radiance fields (NeRFs) achieve high-quality novel-view/novel-pose image synthesis but often require days of training and are extremely slow at in…

Cited by 123SourcePDFScholar
2024

DriveLM: Driving with Graph Visual Question Answering

ECCV 2024oral

"We study how vision-language models (VLMs) trained on web-scale data can be integrated into end-to-end driving systems to boost generalization and enable interactivity with human users. While recent approaches adapt VLMs to driving via single-round visual question answering (VQA), human drivers rea…

2024

Efficient Depth-Guided Urban View Synthesis

ECCV 2024poster

"Recent advances in implicit scene representation enable high-fidelity street view novel view synthesis. However, existing methods optimize a neural radiance field for each scene, relying heavily on dense training images and extensive computation resources. To mitigate this shortcoming, we introduce…

Cited by 1SourcePDFScholar
2024

Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking

ICRA 2024poster

Bin picking is an important building block for many robotic systems, in logistics, production or in household use-cases. In recent years, machine learning methods for the prediction of 6-DoF grasps on diverse and unknown objects have shown promising progress. However, existing approaches only consid…

Cited by 4SourceScholar
2024

GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers

ICLR 2024poster

As transformers are equivariant to the permutation of input tokens, encoding the positional information of tokens is necessary for many tasks. However, since existing positional encoding schemes have been initially designed for NLP tasks, their suitability for vision tasks, which typically exhibit d…

2024

Generalized Predictive Model for Autonomous Driving

CVPR 2024highlight

In this paper we introduce the first large-scale video prediction model in the autonomous driving discipline. To eliminate the restriction of high-cost data collection and empower the generalization ability of our model we acquire massive data from the web and pair it with diverse and high-quality t…

Cited by 61SourcePDFScholar
2024

GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs

CVPR 2024poster

As pretrained text-to-image diffusion models become increasingly powerful recent efforts have been made to distill knowledge from these text-to-image pretrained models for optimizing a text-guided 3D model. Most of the existing methods generate a holistic 3D model from a plain text input. This can b…

Cited by 52SourcePDFScholar
2024

HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting

CVPR 2024poster

Holistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel view synthesis parsing semantic labels and tracking moving objects. Despite considerable progress existing approaches often…

2024

IntrinsicAvatar: Physically Based Inverse Rendering of Dynamic Humans from Monocular Videos via Explicit Ray Tracing

CVPR 2024poster

We present IntrinsicAvatar a novel approach to recovering the intrinsic properties of clothed human avatars including geometry albedo material and environment lighting from only monocular videos. Recent advancements in human-based neural rendering have enabled high-quality geometry and appearance re…

Cited by 12SourcePDFScholar
2024

LISO: Lidar-only Self-Supervised 3D Object Detection

ECCV 2024poster

"3D object detection is one of the most important components in any Self-Driving stack, but current object detectors require costly & slow manual annotation of 3D bounding boxes to perform well. Recently, several methods emerged to generate without human supervision, however, all of these methods ha…

2024

LaRa: Efficient Large-Baseline Radiance Fields

ECCV 2024poster

"Radiance field methods have achieved photorealistic novel view synthesis and geometry reconstruction. But they are mostly applied in per-scene optimization or small-baseline settings. While several recent works investigate feed-forward reconstruction with large baselines by utilizing transformers,…

2024

MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

ECCV 2024oral

"We introduce , an efficient model that, given sparse multi-view images as input, predicts clean feed-forward 3D Gaussians. To accurately localize the Gaussian centers, we build a cost volume representation via plane sweeping, where the cross-view feature similarities stored in the cost volume can p…

2024

Mip-Splatting: Alias-free 3D Gaussian Splatting

CVPR 2024poster

Recently 3D Gaussian Splatting has demonstrated impressive novel view synthesis results reaching high fidelity and efficiency. However strong artifacts can be observed when changing the sampling rate e.g. by changing focal length or camera distance. We find that the source for this phenomenon can be…

Cited by 345SourcePDFScholar
2024

MuRF: Multi-Baseline Radiance Fields

CVPR 2024poster

We present Multi-Baseline Radiance Fields (MuRF) a general feed-forward approach to solving sparse view synthesis under multiple different baseline settings (small and large baselines and different number of input views). To render a target novel view we discretize the 3D space into planes parallel…

2024

NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking

NeurIPS 2024poster

Benchmarking vision-based driving policies is challenging. On one hand, open-loop evaluation with real data is easy, but these results do not reflect closed-loop performance. On the other, closed-loop evaluation is possible in simulation, but is hard to scale due to its significant computational dem…

2024

NeLF-Pro: Neural Light Field Probes for Multi-Scale Novel View Synthesis

CVPR 2024poster

We present NeLF-Pro a novel representation to model and reconstruct light fields in diverse natural scenes that vary in extent and spatial granularity. In contrast to previous fast reconstruction methods that represent the 3D scene globally we model the light field of a scene as a set of local light…

2024

Renovating Names in Open-Vocabulary Segmentation Benchmarks

NeurIPS 2024poster

Names are essential to both human cognition and vision-language models. Open-vocabulary models utilize class names as text prompts to generalize to categories unseen during training. However, the precision of these names is often overlooked in existing datasets. In this paper, we address this undere…

2024

SLEDGE: Synthesizing Driving Environments with Generative Models and Rule-Based Traffic

ECCV 2024poster

"SLEDGE is the first generative simulator for vehicle motion planning trained on real-world driving logs. Its core component is a learned model that is able to generate agent bounding boxes and lane graphs. The model’s outputs serve as an initial state for rule-based traffic simulation. The unique p…

2024

Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

NeurIPS 2024poster

World models can foresee the outcomes of different actions, which is of paramount importance for autonomous driving. Nevertheless, existing driving world models still have limitations in generalization to unseen environments, prediction fidelity of critical details, and action controllability for fl…

2024

WildFusion: Learning 3D-Aware Latent Diffusion Models in View Space

ICLR 2024poster

Modern learning-based approaches to 3D-aware image synthesis achieve high photorealism and 3D-consistent viewpoint changes for the generated images. Existing approaches represent instances in a shared canonical space. However, for in-the-wild datasets a shared canonical system can be difficult to de…

Cited by 6SourcePDFScholar
2023

AG3D: Learning to Generate 3D Avatars from 2D Image Collections

ICCV 2023poster

While progress in 2D generative models of human appearance has been rapid, many applications require 3D avatars that can be animated and rendered. Unfortunately, most existing methods for learning generative models of 3D humans with diverse shape and appearance require 3D training data, which is lim…

Cited by 60PDFScholar
2023

GOOD: Exploring geometric cues for detecting objects in an open world

ICLR 2023poster

We address the task of open-world class-agnostic object detection, i.e., detecting every object in an image by learning from a limited number of base object classes. State-of-the-art RGB-based models suffer from overfitting the training classes and often fail at detecting novel-looking objects. This…

2023

Parting with Misconceptions about Learning-based Vehicle Motion Planning

CoRL 2023poster

The release of nuPlan marks a new era in vehicle motion planning research, offering the first large-scale real-world dataset and evaluation schemes requiring both precise short-term planning and long-horizon ego-forecasting. Existing systems struggle to simultaneously meet both requirements. Indeed,…

Cited by 141SourceScholar
2023

StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis

ICML 2023oral

Text-to-image synthesis has recently seen significant progress thanks to large pretrained language models, large-scale training data, and the introduction of scalable model families such as diffusion and autoregressive models. However, the best-performing models require iterative evaluation to gener…

2022

ARAH: Animatable Volume Rendering of Articulated Human SDFs

ECCV 2022poster

"Combining human body models with differentiable rendering has recently enabled animatable avatars of clothed humans from sparse sets of multi-view RGB videos. While state-of-the-art approaches achieve a realistic appearance with neural radiance fields (NeRF), the inferred geometry often lacks detai…

Cited by 150SourcePDFScholar
2022

KING: Generating Safety-Critical Driving Scenarios for Robust Imitation via Kinematics Gradients

ECCV 2022poster

"Simulators offer the possibility of safe, low-cost development of self-driving systems. However, current driving simulators exhibit naïve behavior models for background traffic. Hand-tuned scenarios are typically added during simulation to induce safety-critical situations. An alternative approach…

2022

MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction

NeurIPS 2022accept

In recent years, neural implicit surface reconstruction methods have become popular for multi-view 3D reconstruction. In contrast to traditional multi-view stereo methods, these approaches tend to produce smoother and more complete reconstructions due to the inductive smoothness bias of neural netwo…

Cited by 505SourcePDFScholar
2022

PINA: Learning a Personalized Implicit Neural Avatar From a Single RGB-D Video Sequence

CVPR 2022poster

We present a novel method to learn Personalized Implicit Neural Avatars (PINA) from a short RGB-D sequence. This allows non-expert users to create a detailed and personalized virtual copy of themselves, which can be animated with realistic clothing deformations. PINA does not require complete scans,…

Cited by 73PDFScholar
2022

PlanT: Explainable Planning Transformers via Object-Level Representations

CoRL 2022poster

Planning an optimal route in a complex environment requires efficient reasoning about the surrounding scene. While human drivers prioritize important objects and ignore details not relevant to the decision, learning-based planners typically extract features from dense, high-dimensional grid represen…

Cited by 114SourcecodeScholar
2022

RegNeRF: Regularizing Neural Radiance Fields for View Synthesis From Sparse Inputs

CVPR 2022oral

Neural Radiance Fields (NeRF) have emerged as a powerful representation for the task of novel view synthesis due to their simplicity and state-of-the-art performance. Though NeRF can produce photorealistic renderings of unseen viewpoints when many input views are available, its performance drops sig…

Cited by 677PDFcodeScholar
2022

VoxGRAF: Fast 3D-Aware Image Synthesis with Sparse Voxel Grids

NeurIPS 2022accept

State-of-the-art 3D-aware generative models rely on coordinate-based MLPs to parameterize 3D radiance fields. While demonstrating impressive results, querying an MLP for every sample along each ray leads to slow rendering. Therefore, existing approaches often render low-resolution feature maps and p…

2022

gDNA: Towards Generative Detailed Neural Avatars

CVPR 2022poster

To make 3D human avatars widely available, we must be able to generate a variety of 3D virtual humans with varied identities and shapes in arbitrary poses. This task is challenging due to the diversity of clothed body shapes, their complex articulations, and the resulting rich, yet stochastic geomet…

Cited by 84PDFScholar
2021

ATISS: Autoregressive Transformers for Indoor Scene Synthesis

NeurIPS 2021poster

The ability to synthesize realistic and diverse indoor furniture layouts automatically or based on partial input, unlocks many applications, from better interactive 3D tools to data synthesis for training and simulation. In this paper, we present ATISS, a novel autoregressive transformer architectur…

2021

KiloNeRF: Speeding Up Neural Radiance Fields With Thousands of Tiny MLPs

ICCV 2021poster

NeRF synthesizes novel views of a scene with unprecedented quality by fitting a neural radiance field to RGB images. However, NeRF requires querying a deep Multi-Layer Perceptron (MLP) millions of times, leading to slow rendering times, even on modern GPUs. In this paper, we demonstrate that real-ti…

Cited by 855PDFcodeScholar
2021

MetaAvatar: Learning Animatable Clothed Human Models from Few Depth Images

NeurIPS 2021poster

In this paper, we aim to create generalizable and controllable neural signed distance fields (SDFs) that represent clothed humans from monocular depth observations. Recent advances in deep learning, especially neural implicit representations, have enabled human shape reconstruction and controllable…

2021

Neural Parts: Learning Expressive 3D Shape Abstractions With Invertible Neural Networks

CVPR 2021poster

Impressive progress in 3D shape extraction led to representations that can capture object geometries with high fidelity. In parallel, primitive-based methods seek to represent objects as semantically consistent part arrangements. However, due to the simplicity of existing primitive representations,…

Cited by 123PDFcodeScholar
2021

SLIM: Self-Supervised LiDAR Scene Flow and Motion Segmentation

ICCV 2021poster

Recently, several frameworks for self-supervised learning of 3D scene flow on point clouds have emerged. Scene flow inherently separates every scene into multiple moving agents and a large class of points following a single rigid sensor motion. However, existing methods do not leverage this property…

Cited by 113PDFScholar
2021

SNARF: Differentiable Forward Skinning for Animating Non-Rigid Neural Implicit Shapes

ICCV 2021poster

Neural implicit surface representations have emerged as a promising paradigm to capture 3D shapes in a continuous and resolution-independent manner. However, adapting them to articulated shapes is non-trivial. Existing approaches learn a backward warp field that maps deformed to canonical points. Ho…

Cited by 258PDFcodeScholar
2021

STEP: Segmenting and Tracking Every Pixel

NeurIPS 2021poster

The task of assigning semantic classes and track identities to every pixel in a video is called video panoptic segmentation. Our work is the first that targets this task in a real-world setting requiring dense interpretation in both spatial and temporal domains. As the ground-truth for this task is…

Cited by 89SourcecodeScholar
2021

Shape As Points: A Differentiable Poisson Solver

NeurIPS 2021oral

In recent years, neural implicit representations gained popularity in 3D reconstruction due to their expressiveness and flexibility. However, the implicit nature of neural implicit representations results in slow inference times and requires careful initialization. In this paper, we revisit the clas…

2021

UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction

ICCV 2021poster

Neural implicit 3D representations have emerged as a powerful paradigm for reconstructing surfaces from multi-view images and synthesizing novel views. Unfortunately, existing methods such as DVR or IDR require accurate per-pixel object masks as supervision. At the same time, neural radiance fields…

Cited by 862PDFcodeScholar
2020

Category Level Object Pose Estimation via Neural Analysis-by-Synthesis

ECCV 2020poster

Many object pose estimation algorithms rely on the analysis-by-synthesis framework which requires explicit representations of individual object instances. In this paper we combine a gradient-based fitting procedure with a parametric neural image synthesis module that is capable of implicitly represe…

Cited by 143SourcePDFScholar
2020

Convolutional Occupancy Networks

ECCV 2020poster

Recently, implicit neural representations have gained popularity for learning-based 3D reconstruction. While demonstrating promising results, most implicit approaches are limited to comparably simple geometry of single objects and do not scale to more complicated or large-scale scenes. The key limit…

2020

Differentiable Volumetric Rendering: Learning Implicit 3D Representations Without 3D Supervision

CVPR 2020poster

Learning-based 3D reconstruction methods have shown impressive results. However, most methods require 3D supervision which is often hard to obtain for real-world datasets. Recently, several works have proposed differentiable rendering techniques to train reconstruction models from RGB images. Unfort…

Cited by 1069PDFcodeScholar
2020

Exploring Data Aggregation in Policy Learning for Vision-Based Urban Autonomous Driving

CVPR 2020poster

Data aggregation techniques can significantly improve vision-based policy learning within a training environment, e.g., learning to drive in a specific simulation condition. However, as on-policy data is sequentially sampled and added in an iterative manner, the policy can specialize and overfit to…

Cited by 104PDFcodeScholar
2020

GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis

NeurIPS 2020poster

While 2D generative adversarial networks have enabled high-resolution image synthesis, they largely lack an understanding of the 3D world and the image formation process. Thus, they do not provide precise control over camera viewpoint or object pose. To address this problem, several recent approache…

2020

Label Efficient Visual Abstractions for Autonomous Driving

IROS 2020poster

It is well known that semantic segmentation can be used as an effective intermediate representation for learning driving policies. However, the task of street scene semantic segmentation requires expensive annotations. Furthermore, segmentation algorithms are often trained irrespective of the actual…

Cited by 50SourceScholar
2020

Learning Unsupervised Hierarchical Part Decomposition of 3D Objects From a Single RGB Image

CVPR 2020poster

Humans perceive the 3D world as a set of distinct objects that are characterized by various low-level (geometry, reflectance) and high-level (connectivity, adjacency, symmetry) properties. Recent methods based on convolutional neural networks (CNNs) demonstrated impressive progress in 3D reconstruct…

Cited by 130PDFcodeScholar
2020

On Joint Estimation of Pose, Geometry and svBRDF From a Handheld Scanner

CVPR 2020poster

We propose a novel formulation for joint recovery of camera pose, object geometry and spatially-varying BRDF. The input to our approach is a sequence of RGB-D images captured by a mobile, hand-held scanner that actively illuminates the scene with point light sources. Compared to previous works that…

Cited by 55PDFScholar
2020

Towards Unsupervised Learning of Generative Models for 3D Controllable Image Synthesis

CVPR 2020poster

In recent years, Generative Adversarial Networks have achieved impressive results in photorealistic image synthesis. This progress nurtures hopes that one day the classical rendering pipeline can be replaced by efficient models that are learned directly from images. However, current image synthesis…

Cited by 180PDFcodeScholar
2019

Connecting the Dots: Learning Representations for Active Monocular Depth Estimation

CVPR 2019poster

We propose a technique for depth estimation with a monocular structured-light camera, i.e., a calibrated stereo set-up with one camera and one laser projector. Instead of formulating the depth estimation via a correspondence search problem, we show that a simple convolutional architecture is suffici…

Cited by 43PDFScholar
2019

MOTS: Multi-Object Tracking and Segmentation

CVPR 2019poster

This paper extends the popular task of multi-object tracking to multi-object tracking and segmentation (MOTS). Towards this goal, we create dense pixel-level annotations for two existing tracking datasets using a semi-automatic annotation procedure. Our new annotations comprise 65,213 pixel masks fo…

Cited by 699PDFScholar
2019

Occupancy Flow: 4D Reconstruction by Learning Particle Dynamics

ICCV 2019poster

Deep learning based 3D reconstruction techniques have recently achieved impressive results. However, while state-of-the-art methods are able to output complex 3D geometry, it is not clear how to extend these results to time-varying topologies. Approaches treating each time step individually lack con…

Cited by 314PDFScholar
2019

Occupancy Networks: Learning 3D Reconstruction in Function Space

CVPR 2019oral

With the advent of deep neural networks, learning-based approaches for 3D reconstruction have gained popularity. However, unlike for images, in 3D there is no canonical representation which is both computationally and memory efficient yet allows for representing high-resolution geometry of arbitrary…

Cited by 3382PDFcodeScholar
2019

PointFlowNet: Learning Representations for Rigid Motion Estimation From Point Clouds

CVPR 2019poster

Despite significant progress in image-based 3D scene flow estimation, the performance of such approaches has not yet reached the fidelity required by many applications. Simultaneously, these applications are often not restricted to image-based estimation: laser scanners provide a popular alternative…

Cited by 140PDFcodeScholar
2019

Project AutoVision: Localization and 3D Scene Perception for an Autonomous Vehicle with a Multi-Camera System

ICRA 2019poster

Project AutoVision aims to develop localization and 3D scene perception capabilities for a self-driving vehicle. Such capabilities will enable autonomous navigation in urban and rural environments, in day and night, and with cameras as the only exteroceptive sensors. The sensor suite employs many ca…

Cited by 149SourceScholar
2019

Real-Time Dense Mapping for Self-Driving Vehicles using Fisheye Cameras

ICRA 2019poster

We present a real-time dense geometric mapping algorithm for large-scale environments. Unlike existing methods which use pinhole cameras, our implementation is based on fisheye cameras whose large field of view benefits various computer vision applications for self-driving vehicles such as visual-in…

Cited by 47SourceScholar
2019

Superquadrics Revisited: Learning 3D Shape Parsing Beyond Cuboids

CVPR 2019poster

Abstracting complex 3D shapes with parsimonious part-based representations has been a long standing goal in computer vision. This paper presents a learning-based solution to this problem which goes beyond the traditional 3D cuboid representation by exploiting superquadrics as atomic elements. We dem…

Cited by 0PDFcodeScholar
2019

Taking a Deeper Look at the Inverse Compositional Algorithm

CVPR 2019oral

In this paper, we provide a modern synthesis of the classic inverse compositional algorithm for dense image alignment. We first discuss the assumptions made by this well-established technique, and subsequently propose to relax these assumptions by incorporating data-driven priors into this model. Mo…

Cited by 63PDFcodeScholar
2019

Texture Fields: Learning Texture Representations in Function Space

ICCV 2019oral

In recent years, substantial progress has been achieved in learning-based reconstruction of 3D objects. At the same time, generative models were proposed that can generate highly realistic images. However, despite this success in these closely related tasks, texture reconstruction of 3D objects has…

Cited by 368PDFcodeScholar
2018

Learning Priors for Semantic 3D Reconstruction

ECCV 2018poster

We present a novel semantic 3D reconstruction framework which embeds variational regularization into a neural network. Our network performs a fixed number of unrolled multi-scale optimization iterations with shared interaction weights. In contrast to existing variational methods for semantic 3D reco…

Cited by 56SourcePDFScholar
2018

RayNet: Learning Volumetric 3D Reconstruction With Ray Potentials

CVPR 2018poster

In this paper, we consider the problem of reconstructing a dense 3D model using images captured from different views. Recent methods based on convolutional neural networks (CNN) allow learning the entire task from data. However, they do not incorporate the physics of image formation such as perspect…

2018

Robust Dense Mapping for Large-Scale Dynamic Environments

ICRA 2018poster

We present a stereo-based dense mapping algorithm for large-scale dynamic urban environments. In contrast to other existing methods, we simultaneously reconstruct the static background, the moving objects, and the potentially moving but currently stationary objects separately, which is desirable for…

Cited by 168SourcecodeScholar
2018

SphereNet: Learning Spherical Representations for Detection and Classification in Omnidirectional Images

ECCV 2018poster

Omnidirectional cameras offer great benefits over classical cameras wherever a wide field of view is essential, such as in virtual reality applications or in autonomous robots. Unfortunately, standard convolutional neural networks are not well suited for this scenario as the natural projection surfa…

2018

Towards Robust Visual Odometry with a Multi-Camera System

IROS 2018poster

We present a visual odometry (VO) algorithm for a multi-camera system and robust operation in challenging environments. Our algorithm consists of a pose tracker and a local mapper. The tracker estimates the current pose by minimizing photometric errors between the most recent keyframe and the curren…

Cited by 56SourceScholar
2018

Unsupervised Learning of Multi-Frame Optical Flow with Occlusions

ECCV 2018poster

Learning optical flow with neural networks is hampered by the need for obtaining training data with associated ground truth. Unsupervised learning is a promising direction, yet the performance of current unsupervised methods is still limited. In particular, the lack of proper occlusion handling in c…

Cited by 232SourcePDFScholar
2017

A Multi-View Stereo Benchmark With High-Resolution Images and Multi-Camera Videos

CVPR 2017poster

Motivated by the limitations of existing multi-view stereo benchmarks, we present a novel dataset for this task. Towards this goal, we recorded a variety of indoor and outdoor scenes using a high-precision laser scanner and captured both high-resolution DSLR imagery as well as synchronized low-resol…

Cited by 1009PDFScholar
2017

Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks

ICML 2017poster

Variational Autoencoders (VAEs) are expressive latent variable models that can be used to learn complex probability distributions from training data. However, the quality of the resulting model crucially relies on the expressiveness of the inference model. We introduce Adversarial Variational Bayes…

Cited by 679SourcePDFScholar
2017

Bounding Boxes, Segmentations and Object Coordinates: How Important Is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios?

ICCV 2017poster

Existing methods for 3D scene flow estimation often fail in the presence of large displacement or local ambiguities, e.g., at texture-less or reflective surfaces. However, these challenges are omnipresent in dynamic road scenes, which is the focus of this work. Our main contribution is to overcome t…

Cited by 189PDFScholar
2017

Direct visual odometry for a fisheye-stereo camera

IROS 2017poster

We present a direct visual odometry algorithm for a fisheye-stereo camera. Our algorithm performs simultaneous camera motion estimation and semi-dense reconstruction. The pipeline consists of two threads: a tracking thread and a mapping thread. In the tracking thread, we estimate the camera pose via…

Cited by 49SourceScholar
2017

Slow Flow: Exploiting High-Speed Cameras for Accurate and Diverse Optical Flow Reference Data

CVPR 2017oral

Existing optical flow datasets are limited in size and variability due to the difficulty of capturing dense ground truth. In this paper, we tackle this problem by tracking pixels through densely sampled space-time volumes recorded with a high-speed video camera. Our model exploits the linearity of s…

Cited by 97PDFScholar
2017

Toroidal Constraints for Two-Point Localization Under High Outlier Ratios

CVPR 2017poster

Localizing a query image against a 3D model at large scale is a hard problem, since 2D-3D matches become more and more ambiguous as the model size increases. This creates a need for pose estimation strategies that can handle very low inlier ratios. In this paper, we draw new insights on the geometri…

Cited by 42PDFScholar
2016

Patches, Planes and Probabilities: A Non-Local Prior for Volumetric 3D Reconstruction

CVPR 2016poster

In this paper, we propose a non-local structured prior for volumetric multi-view 3D reconstruction. Towards this goal, we present a novel Markov random field model based on ray potentials in which assumptions about large 3D surface patches such as planarity or Manhattan world constraints can be effi…

Cited by 42PDFScholar
2016

Semantic Instance Annotation of Street Scenes by 3D to 2D Label Transfer

CVPR 2016poster

This supplementary material provides additional illustrations, visualizations and experiments. We start by showing the color coding and label mapping used for the semantic and instance label results in the paper. Then we provide more details about the 3D fold/curb detection and parameter settings th…

Cited by 220PDFScholar
2015

FollowMe: Efficient Online Min-Cost Flow Tracking With Bounded Memory and Computation

ICCV 2015poster

One of the most popular approaches to multi-target tracking is tracking-by-detection. Current min-cost flow algorithms which solve the data association problem optimally have three main drawbacks: they are computationally expensive, they assume that the whole video is given as a batch, and they sca…

Cited by 132PDFScholar