← Search

Fredo Durand

20 accepted papers

2025

Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences

ICCV 2025poster

Learning alignment between language and vision is a fundamental challenge, especially as multimodal data becomes increasingly detailed and complex. Existing methods often rely on collecting human or AI preferences, which can be costly and time-intensive. We propose an alternative approach that lever…

2025

From Slow Bidirectional to Fast Autoregressive Video Diffusion Models

CVPR 2025poster

Current video diffusion models achieve impressive generation quality but struggle in interactive applications due to bidirectional attention dependencies. The generation of a single frame requires the model to process the entire sequence, including the future. We address this limitation by adapting…

2024

Alchemist: Parametric Control of Material Properties with Diffusion Models

CVPR 2024poster

We propose a method to control material attributes of objects like roughness metallic albedo and transparency in real images. Our method capitalizes on the generative prior of text-to-image models known for photorealism employing a scalar value and instructions to alter low-level material properties…

Cited by 20SourcePDFScholar
2024

Improved Distribution Matching Distillation for Fast Image Synthesis

NeurIPS 2024oral

Recent approaches have shown promises distilling expensive diffusion models into efficient one-step generators. Amongst them, Distribution Matching Distillation (DMD) produces one-step generators that match their teacher in distribution, i.e., the distillation process does not enforce a one-to-one c…

2023

Diffusion with Forward Models: Solving Stochastic Inverse Problems Without Direct Supervision

NeurIPS 2023spotlight

Denoising diffusion models are a powerful type of generative models used to capture complex distributions of real-world signals. However, their applicability is limited to scenarios where training samples are readily available, which is not always the case in real-world applications. For example, in…

Cited by 95SourcePDFScholar
2023

Neural Groundplans: Persistent Neural Scene Representations from a Single Image

ICLR 2023poster

We present a method to map 2D image observations of a scene to a persistent 3D scene representation, enabling novel view synthesis and disentangled representation of the movable and immovable components of the scene. Motivated by the bird’s-eye-view (BEV) representation commonly used in vision and r…

Cited by 13SourcePDFScholar
2021

Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering

NeurIPS 2021spotlight

Inferring representations of 3D scenes from 2D observations is a fundamental problem of computer graphics, computer vision, and artificial intelligence. Emerging 3D-structured neural scene representations are a promising approach to 3D scene understanding. In this work, we propose a novel neural sce…

Cited by 329SourcePDFScholar
2021

Single-Shot Scene Reconstruction

CoRL 2021poster

We introduce a novel scene reconstruction method to infer a fully editable and re-renderable model of a 3D road scene from a single image. We represent movable objects separately from the immovable background, and recover a full 3D model of each distinct object as well as their spatial relations in…

Cited by 18SourceScholar
2020

DiffTaichi: Differentiable Programming for Physical Simulation

ICLR 2020poster

We present DiffTaichi, a new differentiable programming language tailored for building high-performance differentiable physical simulators. Based on an imperative programming language, DiffTaichi generates gradients of simulation steps using source code transformations that preserve arithmetic inten…

Cited by 486SourceScholar
2020

Painting Many Pasts: Synthesizing Time Lapse Videos of Paintings

CVPR 2020poster

We introduce a new video synthesis task: synthesizing time lapse videos depicting how a given painting might have been created. Artists paint using unique combinations of brushes, strokes, and colors. There are often many possible ways to create a given painting. Our goal is to learn to capture this…

Cited by 13PDFScholar
2019

Computational Mirrors: Blind Inverse Light Transport by Deep Matrix Factorization

NeurIPS 2019poster

We recover a video of the motion taking place in a hidden scene by observing changes in indirect illumination in a nearby uncalibrated visible region. We solve this problem by factoring the observed video into a matrix product between the unknown hidden scene video and an unknown light transport mat…

Cited by 56SourcePDFScholar
2019

Data Augmentation Using Learned Transformations for One-Shot Medical Image Segmentation

CVPR 2019oral

Image segmentation is an important task in many medical applications. Methods based on convolutional neural networks attain state-of-the-art accuracy; however, they typically rely on supervised training with large labeled datasets. Labeling medical images requires significant expertise and time, and…

Cited by 608PDFcodeScholar
2019

Visual Deprojection: Probabilistic Recovery of Collapsed Dimensions

ICCV 2019poster

We introduce visual deprojection: the task of recovering an image or video that has been collapsed along a dimension. Projections arise in various contexts, such as long-exposure photography, where a dynamic scene is collapsed in time to produce a motion-blurred image, and corner cameras, where refl…

Cited by 16PDFScholar
2017

Turning Corners Into Cameras: Principles and Methods

ICCV 2017spotlight

We show that walls and other obstructions with edges can be exploited as naturally-occurring "cameras" that reveal the hidden scenes beyond them. In particular, we demonstrate methods for using the subtle spatio-temporal radiance variations that arise on the ground at the base of edges to construct…

Cited by 156PDFScholar
2015

Video Magnification in Presence of Large Motions

CVPR 2015poster

Video magnification reveals subtle variations that would be otherwise invisible to the naked eye. Current techniques require all motion in the video to be very small, which is unfortunately not always the case. Tiny yet meaningful motions are often combined with larger motions, such as the small vib…

Cited by 156SourcePDFScholar
2015

Visual Vibrometry: Estimating Material Properties From Small Motion in Video

CVPR 2015poster

The estimation of material properties is important for scene understanding, with many applications in vision, robotics, and structural engineering. This paper connects fundamentals of vibration mechanics with computer vision techniques in order to infer material properties from small, often impercep…

Cited by 236SourcePDFScholar