← Search

Michael Firman

16 accepted papers

2025

MVSAnywhere: Zero-Shot Multi-View Stereo

CVPR 2025poster

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision.However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs outdoor). Training a general-purpose multi-view stereo model is challenging an…

2024

AirPlanes: Accurate Plane Estimation via 3D-Consistent Embeddings

CVPR 2024poster

Extracting planes from a 3D scene is useful for downstream tasks in robotics and augmented reality. In this paper we tackle the problem of estimating the planar surfaces in a scene from posed images. Our first finding is that a surprisingly competitive baseline results from combining popular cluster…

Cited by 1SourcePDFScholar
2024

DoubleTake: Geometry Guided Depth Estimation

ECCV 2024poster

"Estimating depth from a sequence of posed RGB images is a fundamental computer vision task, with applications in augmented reality, path planning etc. Prior work typically makes use of previous frames in a multi view stereo framework, relying on matching textures in a local neighborhood. In contras…

Cited by 1SourcePDFScholar
2023

Removing Objects From Neural Radiance Fields

CVPR 2023poster

Neural Radiance Fields (NeRFs) are emerging as a ubiquitous scene representation that allows for novel view synthesis. Increasingly, NeRFs will be shareable with other people. Before sharing a NeRF, though, it might be desirable to remove personal information or unsightly objects. Such removal is no…

Cited by 71SourcePDFScholar
2023

Virtual Occlusions Through Implicit Depth

CVPR 2023poster

For augmented reality (AR), it is important that virtual assets appear to 'sit among' real world objects. The virtual element should variously occlude and be occluded by real matter, based on a plausible depth ordering. This occlusion should be consistent over time as the viewer's camera moves. Unfo…

2022

Camera Pose Estimation and Localization with Active Audio Sensing

ECCV 2022poster

"In this work, we show how to estimate a device’s position and orientation indoors by echolocation, i.e., by interpreting the echoes of an audio signal that the device itself emits. Established visual localization methods rely on the device’s camera and yield excellent accuracy if unique visual feat…

2022

SimpleRecon: 3D Reconstruction without 3D Convolutions

ECCV 2022poster

"Traditionally, 3D indoor scene reconstruction from posed images happens in two phases: per-image depth estimation, followed by depth merging and surface reconstruction. Recently, a family of methods have emerged that perform reconstruction directly in final 3D volumetric feature space. While these…

2021

Single Image Depth Prediction With Wavelet Decomposition

CVPR 2021poster

We present a novel method for predicting accurate depths from monocular images with high efficiency. This optimal efficiency is achieved by exploiting wavelet decomposition, which is integrated in a fully differentiable encoder-decoder architecture. We demonstrate that we can reconstruct high-fideli…

Cited by 82PDFcodeScholar
2021

The Temporal Opportunist: Self-Supervised Multi-Frame Monocular Depth

CVPR 2021poster

Self-supervised monocular depth estimation networks are trained to predict scene depth using nearby frames as a supervision signal during training. However, for many applications, sequence information in the form of video frames is also available at test time. The vast majority of monocular networks…

Cited by 346PDFcodeScholar
2020

Footprints and Free Space From a Single Color Image

CVPR 2020oral

Understanding the shape of a scene from a single color image is a formidable computer vision task. However, most methods aim to predict the geometry of surfaces that are visible to the camera, which is of limited use when planning paths for robots or augmented reality agents. Such agents can only mo…

Cited by 24PDFcodeScholar
2020

Learning Stereo from Single Images

ECCV 2020poster

Supervised deep networks are among the best methods for finding correspondences in stereo image pairs. Like all supervised approaches, these networks require ground truth data during training. However, collecting large quantities of accurate dense correspondence data is very challenging. We propose…

2019

Digging Into Self-Supervised Monocular Depth Estimation

ICCV 2019poster

Per-pixel ground-truth depth data is challenging to acquire at scale. To overcome this limitation, self-supervised learning has emerged as a promising alternative for training models to perform monocular depth estimation. In this paper, we propose a set of improvements, which together result in both…

Cited by 2896PDFcodeScholar
2018

DiverseNet: When One Right Answer Is Not Enough

CVPR 2018poster

Many structured prediction tasks in machine vision have a collection of acceptable answers, instead of one definitive ground truth answer. Segmentation of images, for example, is subject to human labeling bias. Similarly, there are multiple possible pixel values that could plausibly complete occlude…

Cited by 38SourcePDFScholar
2016

Structured Prediction of Unobserved Voxels From a Single Depth Image

CVPR 2016oral

Building a complete 3D model of a scene, given only a single depth image, is underconstrained. To gain a full volumetric model, one needs either multiple views, or a single view together with a library of unambiguous 3D models that will fit the shape of each individual object in the scene. We hypot…

Cited by 206PDFScholar