← Search

Felix Heide

66 accepted papers

2026

LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding

AAAI 2026technical

Large-scale scene data is essential for training and testing in robot learning. Neural reconstruction methods have promised the capability of reconstructing large physically-grounded outdoor scenes from captured sensor data. However, these methods have baked-in static environments and only allow for

Cited by 0SourcePDFScholar
2026

TruckDrive: Long-Range Autonomous Highway Driving Dataset

CVPR 2026

Safe highway autonomy for heavy trucks remains an open and unsolved challenge: due to long braking distances, scene understanding of hundreds of meters is required for anticipatory planning and to allow safe braking margins. However, existing driving datasets primarily cover urban scenes, with perce

Cited by 0SourceScholar
2025

A Multi-Modal Benchmark for Long-Range Depth Evaluation in Adverse Weather Conditions

IROS 2025

Depth estimation is a cornerstone computer vision application that is critical for scene understanding and autonomous driving. In real-world scenarios, achieving reliable depth perception under adverse weather—e.g. in fog and rain—is crucial to ensure safety and system robustness. However, quantitat

Cited by 0SourceScholar
2025

Dual Exposure Stereo for Extended Dynamic Range 3D Imaging

CVPR 2025poster

Achieving robust stereo 3D imaging under diverse illumination conditions is an important however challenging task, largely due to the limited dynamic ranges (DRs) of cameras, which are significantly smaller than real world DR. As a result, the accuracy of existing stereo depth estimation methods is…

Cited by 0SourcePDFScholar
2025

Lidar Waveforms are Worth 40x128x33 Words

ICCV 2025poster

Lidar has become crucial for autonomous driving, providing high-resolution 3D scans that are key for accurate scene understanding. To this end, lidar sensors measure the time-resolved full waveforms from the returning laser light, which a subsequent digital signal processor (DSP) converts to point c…

Cited by 0SourcePDFScholar
2025

Neural Atlas Graphs for Dynamic Scene Decomposition and Editing

NeurIPS 2025spotlight

Learning editable high-resolution scene representations for dynamic scenes is an open problem with applications across the domains from autonomous driving to creative editing - the most successful approaches today make a trade-off between editability and supporting scene complexity: neural atlases r…

Cited by 0SourcecodeScholar
2025

Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments

CVPR 2025poster

We introduce Scenario Dreamer, a fully data-driven generative simulator for autonomous vehicle planning that generates both the initial traffic scene--comprising a lane graph and agent bounding boxes--and closed-loop agent behaviours. Existing methods for generating driving simulation environments e…

Cited by 3SourcePDFScholar
2025

Self-Supervised Sparse Sensor Fusion for Long Range Perception

ICCV 2025poster

Outside of urban hubs, autonomous cars and trucks have to master driving on intercity highways. Safe, long-distance highway travel at speeds exceeding 100 km/h demands perception distances of at least 250 m, which is about five times the 50-100m typically addressed in city driving, to allow sufficie…

Cited by 0SourcePDFScholar
2025

Separating the Wheat from the Chaff: Spatio-Temporal Transformer with View-interweaved Attention for Photon-Efficient Depth Sensing

AAAI 2025technical

Time-resolved imaging is an emerging sensing modality that has been shown to enable advanced applications, including remote sensing, fluorescence lifetime imaging, and even non-line-of-sight sensing. Single-photon avalanche diodes (SPADs) outperform relevant time-resolved imaging technologies thanks…

Cited by 0SourcePDFScholar
2024

CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning

CoRL 2024poster

Evaluating autonomous vehicle stacks (AVs) in simulation typically involves replaying driving logs from real-world recorded traffic. However, agents replayed from offline data are not reactive and hard to intuitively control. Existing approaches address these challenges by proposing methods that rel…

Cited by 6SourceScholar
2024

Flow-Guided Online Stereo Rectification for Wide Baseline Stereo

CVPR 2024poster

Stereo rectification is widely considered "solved" due to the abundance of traditional approaches to perform rectification. However autonomous vehicles and robots in-the-wild require constant re-calibration due to exposure to various environmental factors including vibration and structural stress wh…

Cited by 2SourcePDFScholar
2024

Gated Fields: Learning Scene Reconstruction from Gated Videos

CVPR 2024poster

Reconstructing outdoor 3D scenes from temporal observations is a challenge that recent work on neural fields has offered a new avenue for. However existing methods that recover scene properties such as geometry appearance or radiance solely from RGB captures often fail when handling poorly-lit or te…

Cited by 0SourcePDFScholar
2024

Neural Exposure Fusion for High-Dynamic Range Object Detection

CVPR 2024poster

Computer vision in unconstrained outdoor scenarios must tackle challenging high dynamic range (HDR) scenes and rapidly changing illumination conditions. Existing methods address this problem with multi-capture HDR sensors and a hardware image signal processor (ISP) that produces a single fused image…

Cited by 2SourcePDFScholar
2024

Neural Spline Fields for Burst Image Fusion and Layer Separation

CVPR 2024poster

Each photo in an image burst can be considered a sample of a complex 3D scene: the product of parallax diffuse and specular materials scene motion and illuminant variation. While decomposing all of these effects from a stack of misaligned images is a highly ill-conditioned task the conventional alig…

Cited by 13SourcePDFScholar
2024

Polarization Wavefront Lidar: Learning Large Scene Reconstruction from Polarized Wavefronts

CVPR 2024poster

Lidar has become a cornerstone sensing modality for 3D vision especially for large outdoor scenarios and autonomous driving. Conventional lidar sensors are capable of providing centimeter-accurate distance information by emitting laser pulses into a scene and measuring the time-of-flight (ToF) of th…

Cited by 1SourcePDFScholar
2024

Robust Depth Enhancement via Polarization Prompt Fusion Tuning

CVPR 2024poster

Existing depth sensors are imperfect and may provide inaccurate depth values in challenging scenarios such as in the presence of transparent or reflective objects. In this work we present a general framework that leverages polarization imaging to improve inaccurate depth measurements from various de…

2024

SAMFusion: Sensor-Adaptive Multimodal Fusion for 3D Object Detection in Adverse Weather

ECCV 2024poster

"Multimodal sensor fusion is an essential capability for autonomous robots, enabling object detection and decision-making in the presence of failing or uncertain inputs. While recent fusion methods excel in normal environmental conditions, these approaches fail in adverse weather, e.g., heavy fog, s…

2024

Spectral and Polarization Vision: Spectro-polarimetric Real-world Dataset

CVPR 2024highlight

Image datasets are essential not only in validating existing methods in computer vision but also in developing new methods. Many image datasets exist consisting of trichromatic intensity images taken with RGB cameras which are designed to replicate human vision. However polarization and spectrum the…

Cited by 3SourcePDFScholar
2023

Gated Stereo: Joint Depth Estimation From Gated and Wide-Baseline Active Stereo Cues

CVPR 2023highlight

We propose Gated Stereo, a high-resolution and long-range depth estimation technique that operates on active gated stereo images. Using active and high dynamic range passive captures, Gated Stereo exploits multi-view cues alongside time-of-flight intensity cues from active gating. To this end, we pr…

Cited by 8SourcePDFScholar
2023

Kissing to Find a Match: Efficient Low-Rank Permutation Representation

NeurIPS 2023poster

Permutation matrices play a key role in matching and assignment problems across the fields, especially in computer vision and robotics. However, memory for explicitly representing permutation matrices grows quadratically with the size of the problem, prohibiting large problem instances. In this work…

Cited by 3SourcePDFScholar
2023

LiDAR-in-the-Loop Hyperparameter Optimization

CVPR 2023poster

LiDAR has become a cornerstone sensing modality for 3D vision. LiDAR systems emit pulses of light into the scene, take measurements of the returned signal, and rely on hardware digital signal processing (DSP) pipelines to construct 3D point clouds from these measurements. The resulting point clouds…

Cited by 3SourcePDFScholar
2023

Multi-view Spectral Polarization Propagation for Video Glass Segmentation

ICCV 2023poster

In this paper, we present the first polarization-guided video glass segmentation propagation solution (PGVS-Net) that can robustly and coherently propagate glass segmentation in RGB-P video sequences. By leveraging spatiotemporal polarization and color information, our method combines multi-view pol…

Cited by 8PDFScholar
2023

ScatterNeRF: Seeing Through Fog with Physically-Based Inverse Neural Rendering

ICCV 2023poster

Vision in adverse weather conditions, whether it be snow, rain, or fog is challenging. In these scenarios, scattering and attenuation severly degrades image quality. Handling such inclement weather conditions, however, is essential to operate autonomous vehicles, drones and robotic applications wher…

Cited by 21PDFScholar
2023

Seeing With Sound: Long-range Acoustic Beamforming for Multimodal Scene Understanding

CVPR 2023poster

Existing autonomous vehicles primarily use sensors that rely on electromagnetic waves which are undisturbed in good environmental conditions but can suffer in adverse scenarios, such as low light or for objects with low reflectance. Moreover, only objects in direct line-of-sight are typically detect…

Cited by 7SourcePDFScholar
2023

Shakes on a Plane: Unsupervised Depth Estimation From Unstabilized Photography

CVPR 2023poster

Modern mobile burst photography pipelines capture and merge a short sequence of frames to recover an enhanced image, but often disregard the 3D nature of the scene they capture, treating pixel motion between images as a 2D aggregation problem. We show that in a "long-burst", forty-two 12-megapixel R…

Cited by 11SourcePDFScholar
2023

Single Depth-image 3D Reflection Symmetry and Shape Prediction

ICCV 2023poster

In this paper, we present Iterative Symmetry Completion Network (ISCNet), a single depth-image shape completion method that exploits reflective symmetry cues to obtain more detailed shapes. The efficacy of single depth-image shape completion methods is often sensitive to the accuracy of the symmetry…

Cited by 7PDFScholar
2023

The Differentiable Lens: Compound Lens Search Over Glass Surfaces and Materials for Object Detection

CVPR 2023poster

Most camera lens systems are designed in isolation, separately from downstream computer vision methods. Recently, joint optimization approaches that design lenses alongside other components of the image acquisition and processing pipeline--notably, downstream neural networks--have achieved improved…

2022

All You Need Is RAW: Defending against Adversarial Attacks with Camera Image Pipelines

ECCV 2022poster

"Existing neural networks for computer vision tasks are vulnerable to adversarial attacks: adding imperceptible perturbations to the input images can fool these models to make a false prediction on an image that was correctly predicted without the perturbation. Various defense methods have proposed…

Cited by 11SourcePDFScholar
2022

Biologically Inspired Dynamic Thresholds for Spiking Neural Networks

NeurIPS 2022accept

The dynamic membrane potential threshold, as one of the essential properties of a biological neuron, is a spontaneous regulation mechanism that maintains neuronal homeostasis, i.e., the constant overall spiking firing rate of a neuron. As such, the neuron firing rate is regulated by a dynamic spikin…

Cited by 35SourcePDFScholar
2022

Gated2Gated: Self-Supervised Depth Estimation From Gated Images

CVPR 2022oral

Gated cameras hold promise as an alternative to scanning LiDAR sensors with high-resolution 3D depth that is robust to back-scatter in fog, snow, and rain. Instead of sequentially scanning a scene and directly recording depth via the photon time-of-flight, as in pulsed LiDAR sensors, gated imagers e…

Cited by 20PDFcodeScholar
2022

GenSDF: Two-Stage Learning of Generalizable Signed Distance Functions

NeurIPS 2022accept

We investigate the generalization capabilities of neural signed distance functions (SDFs) for learning 3D object representations for unseen and unlabeled point clouds. Existing methods can fit SDFs to a handful of object classes and boast fine detail or fast inference speeds, but do not generalize w…

2022

Glass Segmentation Using Intensity and Spectral Polarization Cues

CVPR 2022poster

Transparent and semi-transparent materials pose significant challenges for existing scene understanding and segmentation algorithms due to their lack of RGB texture which impedes the extraction of meaningful features. In this work, we exploit that the light-matter interactions on glass materials pro…

Cited by 93PDFScholar
2022

Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction

ICLR 2022spotlight

Robust multi-agent trajectory prediction is essential for the safe control of robotic systems. A major challenge is to efficiently learn a representation that approximates the true joint distribution of contextual, social, and temporal information to enable planning. We propose Latent Variable Seque…

2022

LiDAR Snowfall Simulation for Robust 3D Object Detection

CVPR 2022oral

3D object detection is a central task for applications such as autonomous driving, in which the system needs to localize and classify surrounding traffic agents, even in the presence of adverse weather. In this paper, we address the problem of LiDAR-based 3D object detection under snowfall. Due to t…

Cited by 147PDFcodeScholar
2022

Spiking Transformers for Event-Based Single Object Tracking

CVPR 2022poster

Event-based cameras bring a unique capability to tracking, being able to function in challenging real-world conditions as a direct result of their high temporal resolution and high dynamic range. These imagers capture events asynchronously that encode rich temporal and spatial information. However,…

Cited by 190PDFScholar
2022

The Implicit Values of a Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement

CVPR 2022oral

Modern smartphones can continuously stream multi-megapixel RGB images at 60Hz, synchronized with high-quality 3D pose information and low-resolution LiDAR-driven depth estimates. During a snapshot photograph, the natural unsteadiness of the photographer's hands offers millimeter-scale variation in c…

Cited by 17PDFcodeScholar
2021

End-to-End High Dynamic Range Camera Pipeline Optimization

CVPR 2021poster

With a 280 dB dynamic range, the real world is a High Dynamic Range (HDR) world. Today's sensors cannot record this dynamic range in a single shot. Instead, HDR cameras acquire multiple measurements with different exposures, gains and photodiodes, from which an Image Signal Processor (ISP) reconstru…

Cited by 28PDFScholar
2021

Gated3D: Monocular 3D Object Detection From Temporal Illumination Cues

ICCV 2021poster

Today's state-of-the-art methods for 3D object detection are based on lidar, stereo, or monocular cameras. Lidar-based methods achieve the best accuracy, but have a large footprint, high cost, and mechanically-limited angular sampling rates, resulting in low spatial resolution at long ranges. Recent…

Cited by 14PDFScholar
2021

Mask-ToF: Learning Microlens Masks for Flying Pixel Correction in Time-of-Flight Imaging

CVPR 2021poster

We introduce Mask-ToF, a method to reduce flying pixels (FP) in time-of-flight (ToF) depth captures. FPs are pervasive artifacts which occur around depth edges, where light paths from both an object and its background are integrated over the aperture. This light mixes at a sensor pixel to produce er…

Cited by 23PDFcodeScholar
2021

ZeroScatter: Domain Transfer for Long Distance Imaging and Vision Through Scattering Media

CVPR 2021poster

Adverse weather conditions, including snow, rain, and fog, pose a major challenge for both human and computer vision. Handling these environmental conditions is essential for safe decision making, especially in autonomous vehicles, robotics, and drones. Most of today's supervised imaging and vision…

Cited by 15PDFcodeScholar
2020

Hardware-in-the-Loop End-to-End Optimization of Camera Image Processing Pipelines

CVPR 2020oral

Commodity imaging systems rely on hardware image signal processing (ISP) pipelines. These low-level pipelines consist of a sequence of processing blocks that, depending on their hyperparameters, reconstruct a color image from RAW sensor measurements. Hardware ISP hyperparameters have a complex inter…

Cited by 80PDFScholar
2020

Learning Rank-1 Diffractive Optics for Single-Shot High Dynamic Range Imaging

CVPR 2020oral

High-dynamic range (HDR) imaging is an essential imaging modality for a wide range of applications in uncontrolled environments, including autonomous driving, robotics, and mobile phone cameras. However, existing HDR techniques in commodity devices struggle with dynamic scenes due to multi-shot acqu…

Cited by 127PDFScholar
2020

Seeing Around Street Corners: Non-Line-of-Sight Detection and Tracking In-the-Wild Using Doppler Radar

CVPR 2020poster

Conventional sensor systems record information about directly visible objects, whereas occluded scene components are considered lost in the measurement process. Non-line-of-sight (NLOS) methods try to recover such hidden objects from their indirect reflections - faint signal components, traditionall…

Cited by 163PDFcodeScholar
2020

Seeing Through Fog Without Seeing Fog: Deep Multimodal Sensor Fusion in Unseen Adverse Weather

CVPR 2020poster

The fusion of multimodal sensor streams, such as camera, lidar, and radar measurements, plays a critical role in object detection for autonomous vehicles, which base their decision making on these inputs. While existing methods exploit redundant information in good environmental conditions, they fai…

Cited by 577PDFcodeScholar
2020

Single-Shot Monocular RGB-D Imaging Using Uneven Double Refraction

CVPR 2020oral

Cameras that capture color and depth information have become an essential imaging modality for applications in robotics, autonomous driving, virtual, and augmented reality. Existing RGB-D cameras rely on multiple sensors or active illumination with specialized sensors. In this work, we propose a met…

Cited by 12PDFScholar
2019

DeepVoxels: Learning Persistent 3D Feature Embeddings

CVPR 2019oral

In this work, we address the lack of 3D understanding of generative neural networks by introducing a persistent 3D feature embedding for view synthesis. To this end, we propose DeepVoxels, a learned representation that encodes the view-dependent appearance of a 3D scene without having to explicitly…

Cited by 725PDFScholar
2017

Consensus Convolutional Sparse Coding

ICCV 2017poster

Convolutional sparse coding (CSC) is a promising direction for unsupervised learning in computer vision. In contrast to recent supervised methods, CSC allows for convolutional image representations to be learned that are equally useful for high-level vision tasks and low-level image reconstruction a…

Cited by 50PDFcodeScholar
2017

Reconstructing Transient Images From Single-Photon Sensors

CVPR 2017spotlight

Computer vision algorithms build on 2D images or 3D videos that capture dynamic events at the millisecond time scale. However, capturing and analyzing "transient images" at the picosecond scale---i.e., at one trillion frames per second---reveals unprecedented information about a scene and light tran…

Cited by 143PDFScholar
2016

Material Classification Using Raw Time-Of-Flight Measurements

CVPR 2016poster

We propose a material classification method using raw time-of-flight (ToF) measurements. ToF cameras capture the correlation between a reference signal and the temporal response of material to incident illumination. Such measurements encode unique signatures of the material, i.e. the degree of subsu…

Cited by 72PDFScholar
2015

Defocus Deblurring and Superresolution for Time-of-Flight Depth Cameras

CVPR 2015poster

Continuous-wave time-of-flight (ToF) cameras show great promise as low-cost depth image sensors in mobile applications. However, they also suffer from several challenges, including limited illumination intensity, which mandates the use of large numerical aperture lenses, and thus results in a shallo…

Cited by 37SourcePDFScholar