← Search

Steven M. Seitz

19 accepted papers

2026

MusicInfuser: Making Video Diffusion Listen and Dance

CVPR 2026

We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized with specified music tracks. Rather than training a multimodal audio-video or audio-motion model from scratch, our method demonstrates how existing video d

Cited by 0SourcecodeScholar
2024

Generative Powers of Ten

CVPR 2024highlight

We present a method that uses a text-to-image model to generate consistent content across multiple image scales enabling extreme semantic zooms into a scene e.g. ranging from a wide-angle landscape view of a forest to a macro shot of an insect sitting on one of the tree branches. We achieve this thr…

Cited by 5SourcePDFScholar
2024

Total Selfie: Generating Full-Body Selfies

CVPR 2024highlight

We present a method to generate full-body selfies from photographs originally taken at arms length. Because self-captured photos are typically taken close up they have limited field of view and exaggerated perspective that distorts facial shapes. We instead seek to generate the photo some one else w…

Cited by 4SourcePDFScholar
2021

Nerfies: Deformable Neural Radiance Fields

ICCV 2021poster

We present the first method capable of photorealistically reconstructing deformable scenes using photos/videos captured casually from mobile phones. Our approach augments neural radiance fields (NeRF) by optimizing an additional continuous volumetric deformation field that warps each observed point…

Cited by 1543PDFcodeScholar
2021

Real-Time High-Resolution Background Matting

CVPR 2021poster

We introduce a real-time, high-resolution background replacement technique which operates at 30fps in 4K resolution, and 60fps for HD on a modern GPU. Our technique is based on background matting, where an additional frame of the background is captured and used to inform the alpha matte and the fore…

Cited by 282PDFcodeScholar
2020

Background Matting: The World Is Your Green Screen

CVPR 2020poster

We propose a method for creating a matte - the per-pixel foreground color and alpha - of a person by taking photos or videos in an everyday setting with a handheld camera. Most existing matting methods require a green screen background or a manually created trimap to produce a good matte. Automatic,…

Cited by 241PDFcodeScholar
2020

Reconstructing NBA Players

ECCV 2020poster

Great progress has been made in 3D body pose and shape estimation from single photos. Yet, state-of-the-art results still suffer from errors due to challenging body poses, modeling clothing, and self occlusions. The domain of basketball games is particularly challenging, due to all of these factors.…

Cited by 50SourcePDFScholar
2017

IM2CAD

CVPR 2017spotlight

Given a single photo of a room and a large database of furniture CAD models, our goal is to reconstruct a scene that is as similar as possible to the scene depicted in the photograph, and composed of objects drawn from the database. We present a completely automatic system to address this IM2CAD pro…

Cited by 263PDFScholar
2016

The MegaFace Benchmark: 1 Million Faces for Recognition at Scale

CVPR 2016poster

Recent face recognition experiments on a major benchmark LFW show stunning performance--a number of algorithms achieve near to perfect score, surpassing human recognition rates. In this paper, we advocate evaluations at the million scale (LFW includes only 13K photos of 5K people). To this end, we h…

Cited by 1109PDFScholar
2015

DynamicFusion: Reconstruction and Tracking of Non-Rigid Scenes in Real-Time

CVPR 2015poster

We present the first dense SLAM system capable of reconstructing non-rigidly deforming scenes in real-time, by fusing together RGBD scans captured from commodity sensors. Our DynamicFusion approach reconstructs scene geometry whilst simultaneously estimating a dense volumetric 6D motion field that w…

Cited by 1195SourcePDFScholar