← Search

Ira Kemelmacher-Shlizerman

25 accepted papers

2026

MusicInfuser: Making Video Diffusion Listen and Dance

CVPR 2026

We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized with specified music tracks. Rather than training a multimodal audio-video or audio-motion model from scratch, our method demonstrates how existing video d

Cited by 0SourcecodeScholar
2026

Test-Time Anchoring for Discrete Diffusion Posterior Sampling

ICML 2026poster

While continuous diffusion models have achieved remarkable success, discrete diffusion offers a unified framework for jointly modeling text and images. Beyond unification, discrete diffusion provides faster inference, finer control, and principled training-free guidance, making it well-suited for po…

Cited by 0SourceScholar
2025

Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

ICLR 2025poster

We present a method for generating video sequences with coherent motion between a pair of input keyframes. We adapt a pretrained large-scale image-to-video diffusion model (originally trained to generate videos moving forward in time from a single input image) for keyframe interpolation, i.e., to pr…

Cited by 7SourcePDFScholar
2025

Linearly Constrained Diffusion Implicit Models

NeurIPS 2025poster

We introduce Linearly Constrained Diffusion Implicit Models (CDIM), a fast and accurate approach to solving noisy linear inverse problems using diffusion models. Traditional diffusion-based inverse methods rely on numerous projection steps to enforce measurement consistency in addition to unconditio…

Cited by 0SourceScholar
2025

Perturb-and-Revise: Flexible 3D Editing with Generative Trajectories

CVPR 2025poster

Recent advancements in text-based diffusion models have accelerated progress in 3D reconstruction and text-based 3D editing. Although existing 3D editing methods excel at modifying color, texture, and style, they struggle with extensive geometric or appearance changes, thus limiting their applicatio…

2024

Generative Powers of Ten

CVPR 2024highlight

We present a method that uses a text-to-image model to generate consistent content across multiple image scales enabling extreme semantic zooms into a scene e.g. ranging from a wide-angle landscape view of a forest to a macro shot of an insect sitting on one of the tree branches. We achieve this thr…

Cited by 5SourcePDFScholar
2024

M&M VTO: Multi-Garment Virtual Try-On and Editing

CVPR 2024highlight

We present M&M VTO-a mix and match virtual try-on method that takes as input multiple garment images text description for garment layout and an image of a person. An example input includes: an image of a shirt an image of a pair of pants "rolled sleeves shirt tucked in" and an image of a person. The…

2024

Total Selfie: Generating Full-Body Selfies

CVPR 2024highlight

We present a method to generate full-body selfies from photographs originally taken at arms length. Because self-captured photos are typically taken close up they have limited field of view and exaggerated perspective that distorts facial shapes. We instead seek to generate the photo some one else w…

Cited by 4SourcePDFScholar
2023

DreamPose: Fashion Video Synthesis with Stable Diffusion

ICCV 2023poster

We present DreamPose, a diffusion-based method for generating animated fashion videos from still images. Given an image and a sequence of human body poses, our method synthesizes a video containing both human and fabric motion. To achieve this, we transform a pretrained text-to-image model (Stable D…

Cited by 55PDFScholar
2023

PersonNeRF: Personalized Reconstruction From Photo Collections

CVPR 2023poster

We present PersonNeRF, a method that takes a collection of photos of a subject (e.g., Roger Federer) captured across multiple years with arbitrary body poses and appearances, and enables rendering the subject with arbitrary novel combinations of viewpoint, body pose, and appearance. PersonNeRF build…

Cited by 23SourcePDFScholar
2023

TryOnDiffusion: A Tale of Two UNets

CVPR 2023poster

Given two images depicting a person and a garment worn by another person, our goal is to generate a visualization of how the garment might look on the input person. A key challenge is to synthesize a photorealistic detail-preserving visualization of the garment, while warping the garment to accommod…

Cited by 129SourcePDFScholar
2022

HumanNeRF: Free-Viewpoint Rendering of Moving People From Monocular Video

CVPR 2022oral

We introduce a free-viewpoint rendering method -- HumanNeRF -- that works on a given monocular video of a human performing complex body motions, e.g. a video from YouTube. Our method enables pausing the video at any frame and rendering the subject from arbitrary new camera viewpoints or even a full…

Cited by 550PDFcodeScholar
2022

StyleSDF: High-Resolution 3D-Consistent Image and Geometry Generation

CVPR 2022oral

We introduce a high resolution, 3D-consistent image and shape generation technique which we call StyleSDF. Our method is trained on single view RGB data only, and stands on the shoulders of StyleGAN2 for image generation, while solving two main challenges in 3D-aware GANs: 1) high-resolution, view-c…

Cited by 374PDFcodeScholar
2021

Real-Time High-Resolution Background Matting

CVPR 2021poster

We introduce a real-time, high-resolution background replacement technique which operates at 30fps in 4K resolution, and 60fps for HD on a modern GPU. Our technique is based on background matting, where an additional frame of the background is captured and used to inform the alpha matte and the fore…

Cited by 282PDFcodeScholar
2020

Background Matting: The World Is Your Green Screen

CVPR 2020poster

We propose a method for creating a matte - the per-pixel foreground color and alpha - of a person by taking photos or videos in an everyday setting with a handheld camera. Most existing matting methods require a green screen background or a manually created trimap to produce a good matte. Automatic,…

Cited by 241PDFcodeScholar
2020

Lifespan Age Transformation Synthesis

ECCV 2020poster

We address the problem of single photo age progression and regression---the prediction of how a person might look in the future, or how they looked in the past. Most existing aging methods are limited to changing the texture, overlooking transformations in head shape that occur during the human agin…

Cited by 146SourcePDFScholar
2020

Reconstructing NBA Players

ECCV 2020poster

Great progress has been made in 3D body pose and shape estimation from single photos. Yet, state-of-the-art results still suffer from errors due to challenging body poses, modeling clothing, and self occlusions. The domain of basketball games is particularly challenging, due to all of these factors.…

Cited by 50SourcePDFScholar
2020

The Cone of Silence: Speech Separation by Localization

NeurIPS 2020oral

Given a multi-microphone recording of an unknown number of speakers talking concurrently, we simultaneously localize the sources and separate the individual speakers. At the core of our method is a deep network, in the waveform domain, which isolates sources within an angular region $\theta \pm w/2$…

2016

The MegaFace Benchmark: 1 Million Faces for Recognition at Scale

CVPR 2016poster

Recent face recognition experiments on a major benchmark LFW show stunning performance--a number of algorithms achieve near to perfect score, surpassing human recognition rates. In this paper, we advocate evaluations at the million scale (LFW includes only 13K photos of 5K people). To this end, we h…

Cited by 1109PDFScholar