← Search

Noah Snavely

82 accepted papers

2026

ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild

CVPR 2026

Symmetry detection is a fundamental problem in computer vision, and symmetries serve as powerful priors for downstream tasks. However, existing learning-based methods for detecting 3D symmetries from single images have been almost exclusively trained and evaluated on object-centric or synthetic data

Cited by 0SourcecodeScholar
2026

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

CVPR 2026

The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-grained tasks such as action segmentation, classification, captioning, and retrieval. Furthermore, these benchmarks often rely

Cited by 0SourcecodeScholar
2026

ORBIT: Benchmarking SfM in the Wild with 360deg Video

CVPR 2026

Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes.Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard t

Cited by 0SourceScholar
2026

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

CVPR 2026

Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and \pi^3 have a computational cost that scales quadratically with the number of input images, making them inefficient when applied to large image collections. Sequential-reconstruction

Cited by 0SourcecodeScholar
2025

Beyond the Frame: Generating 360deg Panoramic Videos from Perspective Videos

ICCV 2025poster

360deg videos have emerged as a promising medium to represent our dynamic visual world. Compared to the "tunnel vision" of standard cameras, their borderless field of view offers a more complete perspective of our surroundings. While existing video models excel at producing standard videos, their ab…

Cited by 0SourcePDFScholar
2025

C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction

NeurIPS 2025poster

Geometric models like DUSt3R have shown great advances in understanding the geometry of a scene from pairs of photos. However, they fail when the inputs are from vastly different viewpoints (e.g., aerial vs.\ ground) or modalities (e.g., photos vs.\ abstract drawings) compared to what was observed d…

Cited by 0SourcecodeScholar
2025

Can Generative Video Models Help Pose Estimation?

CVPR 2025highlight

Pairwise pose estimation from images with little or no overlap is an open challenge in computer vision. Existing methods, even those trained on large-scale datasets, struggle in these scenarios due to the lack of identifiable correspondences or visual overlap. Inspired by the human ability to infer…

2025

Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features

CVPR 2025highlight

Accurate 3D reconstruction is frequently hindered by visual aliasing, where visually similar but distinct surfaces (aka, doppelgangers), are incorrectly matched. These spurious matches distort the structure-from-motion (SfM) process, leading to misplaced model elements and reduced accuracy. Prior ef…

Cited by 0SourcePDFScholar
2025

FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution

ICCV 2025poster

A versatile video depth estimation model should be consistent and accurate across frames, produce high-resolution depth maps, and support real-time streaming. We propose a method, FlashDepth, that satisfies all three requirements, performing depth estimation for a 2044x1148 streaming video at 24 FPS…

2025

Generating 3D-Consistent Videos from Unposed Internet Photos

CVPR 2025poster

We address the problem of generating videos from unposed internet photos. A handful of input images serve as keyframes, and our model interpolates between them to simulate a path moving between the cameras. Given random images, a model's ability to capture underlying geometry, recognize scene identi…

Cited by 1SourcePDFScholar
2025

LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias

ICLR 2025oral

We propose the Large View Synthesis Model (LVSM), a novel transformer-based approach for scalable and generalizable novel view synthesis from sparse-view inputs. We introduce two architectures: (1) an encoder-decoder LVSM, which encodes input image tokens into a fixed number of 1D latent tokens, fun…

2025

MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos

CVPR 2025award

We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input videos that feature predominantly static scenes with large…

Cited by 18SourcePDFScholar
2025

MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps

ICCV 2025poster

This paper addresses the challenge of learning semantically and functionally meaningful 3D motion priors from real-world videos, in order to enable prediction of future 3D scene motion from a single input image. We propose a novel pixel-aligned Motion Map (MoMap) representation for 3D scene motion,…

Cited by 0SourcePDFScholar
2025

Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

CVPR 2025poster

Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly supervising methods for recovering 3D motion remains challenging due…

2025

Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images

ICCV 2025poster

We present a system using Multimodal LLMs (MLLMs) to analyze a large database with tens of millions of images captured at different times, with the aim of discovering patterns in temporal changes. Specifically, we aim to capture frequent co-occurring changes ("trends") across a city over a certain p…

Cited by 0SourcePDFScholar
2024

MegaScenes: Scene-Level View Synthesis at Scale

ECCV 2024poster

"Scene-level novel view synthesis (NVS) is fundamental to many vision and graphics applications. Recently, pose-conditioned diffusion models have led to significant progress by extracting 3D information from 2D foundation models, but these methods are limited by the lack of scene-level training data…

2024

NeRFiller: Completing Scenes via Generative 3D Inpainting

CVPR 2024poster

We propose NeRFiller an approach that completes missing portions of a 3D capture via generative 3D inpainting using off-the-shelf 2D visual generative models. Often parts of a captured 3D scene or object are missing due to mesh reconstruction failures or a lack of observations (e.g. contact regions…

Cited by 33SourcePDFScholar
2024

Neural Gaffer: Relighting Any Object via Diffusion

NeurIPS 2024poster

Single-image relighting is a challenging task that involves reasoning about the complex interplay between geometry, materials, and lighting. Many prior methods either support only specific categories of images, such as portraits, or require special capture conditions, like using a flashlight. Altern…

Cited by 14SourcePDFScholar
2024

Physics-Based Interaction with 3D Objects via Video Generation

ECCV 2024oral

"Realistic object interactions are crucial for creating immersive virtual experiences, yet synthesizing realistic 3D object dynamics in response to novel interactions remains a significant challenge. Unlike unconditional or text-conditioned dynamics generation, action-conditioned dynamics requires p…

2024

WonderJourney: Going from Anywhere to Everywhere

CVPR 2024poster

We introduce WonderJourney a modular framework for perpetual 3D scene generation. Unlike prior work on view generation that focuses on a single type of scenes we start at any user-provided location (by a text description or an image) and generate a journey through a long sequence of diverse yet cohe…

Cited by 44SourcePDFScholar
2023

ASIC: Aligning Sparse in-the-wild Image Collections

ICCV 2023oral

We present a method for joint alignment of sparse in-the-wild image collections of an object category. Most prior works assume either ground-truth keypoint annotations or a large dataset of images of a single object category. However, neither of the above assumptions hold true for the long-tail of t…

Cited by 21PDFcodeScholar
2023

Accidental Light Probes

CVPR 2023poster

Recovering lighting in a scene from a single image is a fundamental problem in computer vision. While a mirror ball light probe can capture omnidirectional lighting, light probes are generally unavailable in everyday images. In this work, we study recovering lighting from accidental light probes (AL…

Cited by 15SourcePDFScholar
2023

Doppelgangers: Learning to Disambiguate Images of Similar Structures

ICCV 2023oral

We consider the visual disambiguation task of determining whether a pair of visually similar images depict the same or distinct 3D surfaces (e.g., the same or opposite sides of a symmetric building). Illusory image matches, where two images observe distinct but visually similar 3D surfaces, can be c…

Cited by 37PDFcodeScholar
2023

Neural Scene Chronology

CVPR 2023poster

In this work, we aim to reconstruct a time-varying 3D model, capable of rendering photo-realistic renderings with independent control of viewpoint, illumination, and time, from Internet photos of large-scale landmarks. The core challenges are twofold. First, different types of temporal changes, such…

2023

Omnimatte3D: Associating Objects and Their Effects in Unconstrained Monocular Video

CVPR 2023poster

We propose a method to decompose a video into a background and a set of foreground layers, where the background captures stationary elements while the foreground layers capture moving objects along with their associated effects (e.g. shadows and reflections). Our approach is designed for unconstrain…

Cited by 3SourcePDFScholar
2023

Persistent Nature: A Generative Model of Unbounded 3D Worlds

CVPR 2023poster

Despite increasingly realistic image quality, recent 3D image generative models often operate on 3D volumes of fixed extent with limited camera motions. We investigate the task of unconditionally synthesizing unbounded nature scenes, enabling arbitrarily large camera motion while maintaining a persi…

2023

Tracking Everything Everywhere All at Once

ICCV 2023oral

We present a new test-time optimization method for estimating dense and long-range motion from a video sequence. Prior optical flow or particle video tracking algorithms typically operate within limited temporal windows, struggling to track through occlusions and maintain global consistency of estim…

Cited by 170PDFcodeScholar
2022

3D Moments From Near-Duplicate Photos

CVPR 2022poster

We introduce 3D Moments, a new computational photography effect. As input we take a pair of near-duplicate photos, i.e., photos of moving subjects from similar viewpoints, common in people's photo collections. As output, we produce a video that smoothly interpolates the scene motion from the first p…

Cited by 18PDFcodeScholar
2022

Deformable Sprites for Unsupervised Video Decomposition

CVPR 2022oral

We describe a method to extract persistent elements of a dynamic scene from an input video. We represent each scene element as a Deformable Sprite consisting of three components: 1) a 2D texture image for the entire video, 2) per-frame masks for the element, and 3) non-rigid deformations that map th…

Cited by 78PDFScholar
2022

IRON: Inverse Rendering by Optimizing Neural SDFs and Materials From Photometric Images

CVPR 2022oral

We propose a neural inverse rendering pipeline called IRON that operates on photometric images and outputs high-quality 3D content in the format of triangle meshes and material textures readily deployable in existing graphics pipelines. We propose a neural inverse rendering pipeline called IRON that…

Cited by 116PDFScholar
2022

InfiniteNature-Zero: Learning Perpetual View Generation of Natural Scenes from Single Images

ECCV 2022poster

"We present a method for learning to generate unbounded flythrough videos of natural scenes starting from a single view. This capability is learned from a collection of single photographs, without requiring camera poses or even multiple views of each scene. To achieve this, we propose a novel self-s…

2022

Structure and Motion from Casual Videos

ECCV 2022poster

"Casual videos, such as those captured in daily life using a hand-held cell phone, pose problems for conventional structure-from-motion (SfM) techniques: the camera is often roughly stationary (not much parallax), and a large portion of the video may contain moving objects. Under such conditions, st…

Cited by 42SourcePDFScholar
2022

Unsupervised Semantic Segmentation by Distilling Feature Correspondences

ICLR 2022poster

Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clus…

2021

De-Rendering the World's Revolutionary Artefacts

CVPR 2021poster

Recent works have shown exciting results in unsupervised image de-rendering--learning to decompose 3D shape, appearance, and lighting from single-image collections without explicit supervision. However, many of these assume simplistic material and lighting models. We propose a method, termed RADAR,…

Cited by 35PDFcodeScholar
2021

Extreme Rotation Estimation Using Dense Correlation Volumes

CVPR 2021poster

We present a technique for estimating the relative 3D rotation of an RGB image pair in an extreme setting, where the images have little or no overlap. We observe that, even when images do not overlap, there may be rich hidden cues as to their geometric relationship, such as light source directions,…

Cited by 48PDFcodeScholar
2021

IBRNet: Learning Multi-View Image-Based Rendering

CVPR 2021poster

We present a method that synthesizes novel views of complex scenes by interpolating a sparse set of nearby views. The core of our method is a network architecture that includes a multilayer perceptron and a ray transformer that estimates radiance and volume density at continuous 5D locations (3D spa…

Cited by 956PDFScholar
2021

Infinite Nature: Perpetual View Generation of Natural Scenes From a Single Image

ICCV 2021poster

We introduce the problem of perpetual view generation - long-range generation of novel views corresponding to an arbitrarily long camera trajectory given a single image. This is a challenging problem that goes far beyond the capabilities of current view synthesis methods, which quickly degenerate wh…

Cited by 169PDFcodeScholar
2021

KeypointDeformer: Unsupervised 3D Keypoint Discovery for Shape Control

CVPR 2021poster

We introduce KeypointDeformer, a novel unsupervised method for shape control through automatically discovered 3D keypoints. We cast this as the problem of aligning a source 3D object to a target 3D object from the same object category. Our method analyzes the difference between the shapes of the two…

Cited by 70PDFcodeScholar
2021

Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes

CVPR 2021poster

We present a method to perform novel view and time synthesis of dynamic scenes, requiring only a monocular video with known camera poses as input. To do this, we introduce Neural Scene Flow Fields, a new representation that models the dynamic scene as a time-variant continuous function of appearance…

Cited by 862PDFScholar
2021

PhySG: Inverse Rendering With Spherical Gaussians for Physics-Based Material Editing and Relighting

CVPR 2021poster

We present an end-to-end inverse rendering pipeline that includes a fully differentiable renderer, and can reconstruct geometry, materials, and illumination from scratch from a set of images. Our rendering framework represents specular BRDFs and environmental illumination using mixtures of spherical…

Cited by 365PDFScholar
2021

Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision

ICCV 2021poster

The abundance and richness of Internet photos of landmarks and cities has led to significant progress in 3D vision over the past two decades, including automated 3D reconstructions of the world's landmarks from tourist photos. However, a major source of information available for these 3D-augmented c…

Cited by 20PDFcodeScholar
2021

Who's Waldo? Linking People Across Text and Images

ICCV 2021poster

We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which is predominantly object-based, our new task masks out the names of people in cap…

Cited by 22PDFcodeScholar
2020

An Analysis of SVD for Deep Rotation Estimation

NeurIPS 2020poster

Symmetric orthogonalization via SVD, and closely related procedures, are well-known techniques for projecting matrices onto O(n) or SO(n). These tools have long been used for applications in computer vision, for example optimal 3D alignment problems solved by orthogonal Procrustes, rotation averagin…

2020

DualSDF: Semantic Shape Manipulation Using a Two-Level Representation

CVPR 2020poster

We are seeing a Cambrian explosion of 3D shape representations for use in machine learning. Some representations seek high expressive power in capturing high-resolution detail. Other approaches seek to represent shapes as compositions of simple parts, which are intuitive for people to understand and…

Cited by 130PDFcodeScholar
2020

Hidden Footprints: Learning Contextual Walkability from 3D Human Trails

ECCV 2020poster

Predicting where people can walk in a scene is important for many tasks, including autonomous driving systems and human behavior analysis. Yet learning a computational model for this purpose is challenging due to semantic ambiguity and a lack of labeled data: current datasets only have labels on whe…

2020

Learning Feature Descriptors using Camera Pose Supervision

ECCV 2020poster

Recent research on learned visual descriptors has shown promising improvements in correspondence estimation, a key component of many 3D vision tasks. However, existing descriptor learning frameworks typically require ground-truth correspondences between feature points for training, which are challen…

Cited by 201SourcePDFScholar
2020

Learning Gradient Fields for Shape Generation

ECCV 2020poster

In this work, we propose a novel technique to generate shapes from point cloud data. A point cloud can be viewed as samples from a distribution of 3D points whose density is concentrated near the surface of the shape. Point cloud generation thus amounts to moving randomly sampled points to high-dens…

2020

Learning to Factorize and Relight a City

ECCV 2020poster

We propose a learning-based framework for disentangling outdoor scenes into temporally-varying illumination and permanent scene factors. Inspired by the classic intrinsic image decomposition, our learning signal builds upon two insights: 1) combining the disentangled factors should reconstruct the o…

2020

Lighthouse: Predicting Lighting Volumes for Spatially-Coherent Illumination

CVPR 2020poster

We present a deep learning solution for estimating the incident illumination at any 3D location within a scene from an input narrow-baseline stereo image pair. Previous approaches for predicting global illumination from images either predict just a single illumination for the entire scene, or separa…

Cited by 119PDFcodeScholar
2020

MetaSDF: Meta-Learning Signed Distance Functions

NeurIPS 2020poster

Neural implicit shape representations are an emerging paradigm that offers many potential benefits over conventional discrete representations, including memory efficiency at a high spatial resolution. Generalizing across shapes with such neural implicit representations amounts to learning priors ove…

2020

Multi-Plane Program Induction with 3D Box Priors

NeurIPS 2020poster

We consider two important aspects in understanding and editing images: modeling regular, program-like texture or patterns in 2D planes, and 3D posing of these planes in the scene. Unlike prior work on image-based program synthesis, which assumes the image contains a single visible 2D plane, we prese…

Cited by 15SourcePDFScholar
2020

NBVC: A Benchmark for Depth Estimation from Narrow-Baseline Video Clips

IROS 2020poster

We present a benchmark for online, video-based depth estimation, a problem that is not covered by the current set of benchmarks for evaluating 3D reconstruction, which focus on offline, batch reconstruction. Online depth estimation from video captured by a moving camera is a key enabling technology…

Cited by 2SourceScholar
2019

DeepView: View Synthesis With Learned Gradient Descent

CVPR 2019oral

We present a novel approach to view synthesis using multiplane images (MPIs). Building on recent advances in learned gradient descent, our algorithm generates an MPI from a set of sparse camera viewpoints. The resulting method incorporates occlusion reasoning, improving performance on challenging sc…

Cited by 516PDFScholar
2019

Learning the Depths of Moving People by Watching Frozen People

CVPR 2019oral

We present a method for predicting dense depth in scenarios where both a monocular camera and people in the scene are freely moving. Existing methods for recovering depth for dynamic, non-rigid objects from monocular video impose strong assumptions on the objects' motion and may only recover sparse…

Cited by 276PDFScholar
2019

Pushing the Boundaries of View Extrapolation With Multiplane Images

CVPR 2019oral

We explore the problem of view synthesis from a narrow baseline pair of images, and focus on generating high-quality view extrapolations with plausible disocclusions. Our method builds upon prior work in predicting a multiplane image (MPI), which represents scene content as a set of RGBA planes with…

Cited by 365PDFScholar
2019

TOUCHDOWN: Natural Language Navigation and Spatial Reasoning in Visual Street Environments

CVPR 2019poster

We study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task. We introduce the Touchdown task and dataset, where an agent must first follow navigation instructions in a Street View environment to a goal position, and then guess a location in its…

Cited by 433PDFcodeScholar
2019

UprightNet: Geometry-Aware Camera Orientation Estimation From Single Images

ICCV 2019poster

We introduce UprightNet, a learning-based approach for estimating 2DoF camera orientation from a single RGB image of an indoor scene. Unlike recent methods that leverage deep learning to perform black-box regression from image to orientation parameters, we propose an end-to-end framework that incorp…

Cited by 57PDFScholar
2018

Discovery of Latent 3D Keypoints via End-to-end Geometric Reasoning

NeurIPS 2018oral

This paper presents KeypointNet, an end-to-end geometric reasoning framework to learn an optimal set of category-specific keypoints, along with their detectors to predict 3D keypoints in a single 2D input image. We demonstrate this framework on 3D pose estimation task by proposing a differentiable p…

2017

Deep Feature Interpolation for Image Content Changes

CVPR 2017poster

We propose Deep Feature Interpolation (DFI), a new data- driven baseline for automatic high-resolution image transformation. As the name suggests, DFI relies only on simple linear interpolation of deep convolutional features from pre-trained convnets. We show that despite its simplicity, DFI can per…

Cited by 386PDFcodeScholar
2016

DeepStereo: Learning to Predict New Views From the World's Imagery

CVPR 2016spotlight

Deep networks have recently enjoyed enormous success when applied to recognition and classification problems in computer vision [22, 32], but their use in graphics problems has been limited ([23, 7] are notable recent exceptions). In this work, we present a novel deep architecture that per- forms ne…

Cited by 798PDFScholar
2015

Material Recognition in the Wild With the Materials in Context Database

CVPR 2015poster

Recognizing materials in real-world images is a challenging task. Real-world materials have rich surface texture, geometry, lighting conditions, and clutter, which combine to make the problem particularly difficult. In this paper, we introduce a new, large-scale, open dataset of materials in the wil…

Cited by 697SourcePDFScholar