← Search

Gabriel Brostow

12 accepted papers

2026

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

CVPR 2026

We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level and by satellite. Faithful reconstructions require good camera coverage, but ground imagery is time-consuming and hard to capture at scale for large outdoor s

Cited by 0SourcecodeScholar
2025

MVSAnywhere: Zero-Shot Multi-View Stereo

CVPR 2025poster

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision.However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs outdoor). Training a general-purpose multi-view stereo model is challenging an…

2025

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

ICCV 2025poster

We introduce the task of Language-Guided Object Placement in Real 3D Scenes. Given a 3D reconstructed point-cloud scene, a 3D asset, and a natural-language instruction, the goal is to place the asset so that the instruction is satisfied. The task demands tackling four intertwined challenges: (a) one…

Cited by 0SourcePDFScholar
2024

AirPlanes: Accurate Plane Estimation via 3D-Consistent Embeddings

CVPR 2024poster

Extracting planes from a 3D scene is useful for downstream tasks in robotics and augmented reality. In this paper we tackle the problem of estimating the planar surfaces in a scene from posed images. Our first finding is that a surprisingly competitive baseline results from combining popular cluster…

Cited by 1SourcePDFScholar
2024

DoubleTake: Geometry Guided Depth Estimation

ECCV 2024poster

"Estimating depth from a sequence of posed RGB images is a fundamental computer vision task, with applications in augmented reality, path planning etc. Prior work typically makes use of previous frames in a multi view stereo framework, relying on matching textures in a local neighborhood. In contras…

Cited by 1SourcePDFScholar
2024

GroundUp: Rapid Sketch-Based 3D City Massing

ECCV 2024poster

"We propose GroundUp, the first sketch-based ideation tool for 3D city massing of urban areas. We focus on early-stage urban design, where sketching is a common tool and the design starts from balancing planned building volumes (masses) and open spaces. With Human-Centered AI in mind, we aim to help…

Cited by 1SourcePDFScholar
2024

INQUIRE: A Natural World Text-to-Image Retrieval Benchmark

NeurIPS 2024poster

We introduce INQUIRE, a text-to-image retrieval benchmark designed to challenge multimodal vision-language models on expert-level queries. INQUIRE includes iNaturalist 2024 (iNat24), a new dataset of five million natural world images, along with 250 expert-level retrieval queries. These queries are…

2024

TAPVid-3D: A Benchmark for Tracking Any Point in 3D

NeurIPS 2024poster

We introduce a new benchmark, TAPVid-3D, for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D). While point tracking in two dimensions (TAP-2D) has many benchmarks measuring performance on real-world videos, such as TAPVid-DAVIS, three-dimensional point tracking has none. To this e…

2021

The Temporal Opportunist: Self-Supervised Multi-Frame Monocular Depth

CVPR 2021poster

Self-supervised monocular depth estimation networks are trained to predict scene depth using nearby frames as a supervision signal during training. However, for many applications, sequence information in the form of video frames is also available at test time. The vast majority of monocular networks…

Cited by 346PDFcodeScholar