← Search

David B Lindell

34 accepted papers

2026

Dark3R: Learning Structure from Motion in the Dark

CVPR 2026

We introduce Dark3R, a framework for structure from motion in the dark that operates directly on raw images with signal-to-noise ratios (SNRs) below -4 dB--a regime where conventional feature- and learning-based methods break down. Our key insight is to adapt large-scale 3D foundation models to extr

Cited by 0SourceScholar
2026

Lumosaic: Hyperspectral Video via Active Illumination and Coded-Exposure Pixels

CVPR 2026

We present Lumosaic, a compact active hyperspectral video system designed for real-time capture of dynamic scenes. Our approach combines a programmable narrowband LED array with a coded-exposure-pixel (CEP) camera capable of high-speed pixel-wise exposure, thereby enabling joint encoding of scene in

Cited by 0SourceScholar
2026

Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation

ICLR 2026poster

The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the availability of captured real-world multi-view data, which is not…

Cited by 0SourcecodeScholar
2026

Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Prediction

CVPR 2026

Reconstructing high-fidelity 3D head geometry from images is critical for a wide range of applications, yet existing methods face fundamental limitations. Traditional photogrammetry achieves exceptional detail but requires extensive camera arrays (25--200+ views), substantial computation, and manual

Cited by 0SourceScholar
2026

Velox: Learning Representations of 4D Geometry and Appearance

CVPR 2026

We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input, i.e., an unstructured dynamic point cloud, to construct. Speci

Cited by 0SourceScholar
2025

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

CVPR 2025poster

Numerous works have recently integrated 3D camera control into foundational text-to-video models, but the resulting camera control is often imprecise, and video generation quality suffers. In this work, we analyze camera motion from a first principles perspective, uncovering insights that enable pre…

Cited by 10SourcePDFScholar
2025

CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models

CVPR 2025poster

Reconstructing photorealistic and dynamic portrait avatars from images is essential to many applications including advertising, visual effects, and virtual reality. Depending on the application, avatar reconstruction involves different capture setups and constraints -- for example, visual effects st…

Cited by 2SourcePDFScholar
2025

Efficient Neural Network Encoding for 3D Color Lookup Tables

AAAI 2025technical

3D color lookup tables (LUTs) enable precise color manipulation by mapping input RGB values to specific output RGB values. 3D LUTs are instrumental in various applications, including video editing, in-camera processing, photographic filters, computer graphics, and color processing for displays. Whi…

2025

Neural Inverse Rendering from Propagating Light

CVPR 2025poster

We present the first system for physically based, neural inverse rendering from multi-viewpoint videos of propagating light. Our approach relies on a time-resolved extension of neural radiance caching -- a technique that accelerates inverse rendering by storing infinite-bounce radiance arriving at a…

Cited by 0SourcePDFScholar
2025

Opportunistic Single-Photon Time of Flight

CVPR 2025poster

Scattered light from pulsed lasers is increasingly part of our ambient illumination, as many devices rely on them for active 3D sensing. In this work, we ask: can these "ambient" light signals be detected and leveraged for passive 3D vision? We show that pulsed lasers, despite being weak and fluctua…

Cited by 0SourcePDFScholar
2025

Reconstructing Heterogeneous Biomolecules via Hierarchical Gaussian Mixtures and Part Discovery

NeurIPS 2025poster

Cryo-EM is a transformational paradigm in molecular biology where computational methods are used to infer 3D molecular structure at atomic resolution from extremely noisy 2D electron microscope images. At the forefront of research is how to model the structure when the imaged particles exhibit non-r…

Cited by 0SourceScholar
2025

SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation

ICLR 2025poster

Methods for image-to-video generation have achieved impressive, photo-realistic quality. However, adjusting specific elements in generated videos, such as object motion or camera movement, is often a tedious process of trial and error, e.g., involving re-generating videos with different random seed…

2025

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

ICLR 2025poster

Modern text-to-video synthesis models demonstrate coherent, photorealistic generation of complex videos from a text description. However, most existing models lack fine-grained control over camera movement, which is critical for downstream applications related to content creation, visual effects, an…

Cited by 38SourcePDFScholar
2024

4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling

CVPR 2024poster

Recent breakthroughs in text-to-4D generation rely on pre-trained text-to-image and text-to-video models to generate dynamic 3D scenes. However current text-to-4D methods face a three-way tradeoff between the quality of scene appearance 3D structure and motion. For example text-to-image models and t…

2024

CryoSPIN: Improving Ab-Initio Cryo-EM Reconstruction with Semi-Amortized Pose Inference

NeurIPS 2024poster

Cryo-EM is an increasingly popular method for determining the atomic resolution 3D structure of macromolecular complexes (eg, proteins) from noisy 2D images captured by an electron microscope. The computational task is to reconstruct the 3D density of the particle, along with 3D pose of the particle…

Cited by 1SourcePDFScholar
2024

Flying with Photons: Rendering Novel Views of Propagating Light

ECCV 2024oral

"We present an imaging and neural rendering technique that seeks to synthesize videos of light propagating through a scene from novel, moving camera viewpoints. Our approach relies on a new ultrafast imaging setup to capture a first-of-its kind, multi-viewpoint video dataset with picosecond-level te…

Cited by 5SourcePDFScholar
2024

MoSS: Monocular Shape Sensing for Continuum Robots

RA-L 2024

Continuum robots are promising candidates for interactive tasks in medical and industrial applications due to their unique shape, compliance, and miniaturization capability. Accurate and real-time shape sensing is essential for such tasks yet remains a challenge. Embedded shape sensing has high hard

Cited by 12SourcecodeScholar
2024

SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark Estimation

CVPR 2024poster

Self-supervised landmark estimation is a challenging task that demands the formation of locally distinct feature representations to identify sparse facial landmarks in the absence of annotated data. To tackle this task existing state-of-the-art (SOTA) methods (1) extract coarse features from backbon…

Cited by 1SourcePDFScholar
2024

TC4D: Trajectory-Conditioned Text-to-4D Generation

ECCV 2024poster

"Recent techniques for text-to-4D generation synthesize dynamic 3D scenes using supervision from pre-trained text-to-video models. However, existing representations, such as deformation models or time-dependent neural representations, are limited in the amount of motion they can generate—they cannot…

Cited by 37SourcePDFScholar
2024

TurboSL: Dense Accurate and Fast 3D by Neural Inverse Structured Light

CVPR 2024poster

We show how to turn a noisy and fragile active triangulation technique--three-pattern structured light with a grayscale camera--into a fast and powerful tool for 3D capture: able to output sub-pixel accurate disparities at megapixel resolution along with reflectance normals and a no-reference estima…

Cited by 5SourcePDFScholar
2023

SparsePose: Sparse-View Camera Pose Regression and Refinement

CVPR 2023poster

Camera pose estimation is a key step in standard 3D reconstruction pipelines that operates on a dense set of images of a single object or scene. However, methods for pose estimation often fail when there are only a few images available because they rely on the ability to robustly identify and match…

Cited by 45SourcePDFScholar
2023

Transient Neural Radiance Fields for Lidar View Synthesis and 3D Reconstruction

NeurIPS 2023spotlight

Neural radiance fields (NeRFs) have become a ubiquitous tool for modeling scene appearance and geometry from multiview imagery. Recent work has also begun to explore how to use additional supervision from lidar or depth sensor measurements in the NeRF framework. However, previous lidar-supervised Ne…

Cited by 21SourcePDFScholar
2022

BACON: Band-Limited Coordinate Networks for Multiscale Scene Representation

CVPR 2022oral

Coordinate-based networks have emerged as a powerful tool for 3D representation and scene reconstruction. These networks are trained to map continuous input coordinates to the value of a signal at each point. Still, current architectures are black boxes: their spectral characteristics cannot be easi…

Cited by 175PDFcodeScholar
2022

Generative Neural Articulated Radiance Fields

NeurIPS 2022accept

Unsupervised learning of 3D-aware generative adversarial networks (GANs) using only collections of single-view 2D photographs has very recently made much progress. These 3D GANs, however, have not been demonstrated for human bodies and the generated radiance fields of existing frameworks are not dir…

Cited by 119SourcePDFScholar
2022

Learning to Solve PDE-constrained Inverse Problems with Graph Networks

ICML 2022spotlight

Learned graph neural networks (GNNs) have recently been established as fast and accurate alternatives for principled solvers in simulating the dynamics of physical systems. In many application domains across science and engineering, however, we are not only interested in a forward simulation but als…

Cited by 45SourcePDFScholar
2022

Residual Multiplicative Filter Networks for Multiscale Reconstruction

NeurIPS 2022accept

Coordinate networks like Multiplicative Filter Networks (MFNs) and BACON offer some control over the frequency spectrum used to represent continuous signals such as images or 3D volumes. Yet, they are not readily applicable to problems for which coarse-to-fine estimation is required, including vario…

2021

AutoInt: Automatic Integration for Fast Neural Volume Rendering

CVPR 2021poster

Numerical integration is a foundational technique in scientific computing and is at the core of many computer vision applications. Among these applications, neural volume rendering has recently been proposed as a new paradigm for view synthesis, achieving photorealistic image quality. However, a fun…

Cited by 281PDFcodeScholar
2020

Disambiguating Monocular Depth Estimation with a Single Transient

ECCV 2020poster

Monocular depth estimation algorithms successfully predict the relative depth order of objects in a scene. However, because of the fundamental scale ambiguity associated with monocular images, these algorithms fail at correctly predicting true metric depth. In this work, we demonstrate how a depth h…

Cited by 35SourcePDFScholar
2020

Non-Line-of-Sight Surface Reconstruction Using the Directional Light-Cone Transform

CVPR 2020oral

We propose a joint albedo-normal approach to non-line-of-sight (NLOS) surface reconstruction using the directional light-cone transform (D-LCT). While current NLOS imaging methods reconstruct either the albedo or surface normals of the hidden scene, the two quantities provide complementary informati…

Cited by 83PDFScholar
2017

Reconstructing Transient Images From Single-Photon Sensors

CVPR 2017spotlight

Computer vision algorithms build on 2D images or 3D videos that capture dynamic events at the millisecond time scale. However, capturing and analyzing "transient images" at the picosecond scale---i.e., at one trillion frames per second---reveals unprecedented information about a scene and light tran…

Cited by 143PDFScholar