← Search

Michaël Gharbi

13 accepted papers

2026

PromptRL: Prompt Matters in RL for Flow-Based Image Generation

ICML 2026poster

Flow matching models (FMs) have revolutionized text-to-image (T2I) generation, with reinforcement learning (RL) serving as a critical post-training strategy for alignment with reward objectives. In this research, we show that current RL pipelines for FMs suffer from two underappreciated yet importan…

Cited by 0SourceScholar
2024

Editable Image Elements for Controllable Synthesis

ECCV 2024poster

"Diffusion models have made significant advances in text-guided synthesis tasks. However, editing user-provided images remains challenging, as the high dimensional noise input space of diffusion models is not naturally suited for image inversion or spatial editing. In this work, we propose an image…

Cited by 8SourcePDFScholar
2024

Improved Distribution Matching Distillation for Fast Image Synthesis

NeurIPS 2024oral

Recent approaches have shown promises distilling expensive diffusion models into efficient one-step generators. Amongst them, Distribution Matching Distillation (DMD) produces one-step generators that match their teacher in distribution, i.e., the distillation process does not enforce a one-to-one c…

2024

Lazy Diffusion Transformer for Interactive Image Editing

ECCV 2024poster

"We introduce a novel diffusion transformer, , that generates partial image updates efficiently. Our approach targets interactive image editing applications in which, starting from a blank canvas or an image, a user specifies a sequence of localized image modifications using binary masks and text pr…

Cited by 8SourcePDFScholar
2024

Learning Subject-Aware Cropping by Outpainting Professional Photos

AAAI 2024technical

How to frame (or crop) a photo often depends on the image subject and its context; e.g., a human portrait. Recent works have defined the subject-aware image cropping task as a nuanced and practical version of image cropping. We propose a weakly-supervised approach (GenCrop) to learn what makes a hig…

Cited by 2SourcePDFScholar
2024

One-step Diffusion with Distribution Matching Distillation

CVPR 2024poster

Diffusion models generate high-quality images but require dozens of forward passes. We introduce Distribution Matching Distillation (DMD) a procedure to transform a diffusion model into a one-step image generator with minimal impact on image quality. We enforce the one-step image generator match the…

Cited by 946SourcePDFScholar
2023

Domain Expansion of Image Generators

CVPR 2023poster

Can one inject new concepts into an already trained generative model, while respecting its existing structure and knowledge? We propose a new task -- domain expansion -- to address this. Given a pretrained generator and novel (but related) domains, we expand the generator to jointly model all domain…

Cited by 17SourcePDFScholar
2023

Semi-Supervised Parametric Real-World Image Harmonization

CVPR 2023poster

Learning-based image harmonization techniques are usually trained to undo synthetic global transformations, applied to a masked foreground in a single ground truth photo. This simulated data does not model many important appearance mismatches (illumination, object boundaries, etc.) between foregroun…

2022

"Spotting Temporally Precise, Fine-Grained Events in Video"

ECCV 2022poster

"We introduce the task of spotting temporally precise, fine-grained events in video (detecting the precise moment in time events occur). Precise spotting requires models to reason globally about the full-time scale of actions and locally to identify subtle frame-to-frame appearance and motion differ…

2022

Any-Resolution Training for High-Resolution Image Synthesis

ECCV 2022poster

"Generative models operate at fixed resolution, even though natural images come in a variety of sizes. As high-resolution details are downsampled away and low-resolution images are discarded altogether, precious supervision is lost. We argue that every pixel matters and create datasets with variable…

2021

Modulated Periodic Activations for Generalizable Local Functional Representations

ICCV 2021poster

Multi-Layer Perceptrons (MLPs) make powerful functional representations for sampling and reconstruction problems involving low-dimensional signals like images,shapes and light fields. Recent works have significantly improved their ability to represent high-frequency content by using periodic activat…

Cited by 162PDFScholar
2021

Video Pose Distillation for Few-Shot, Fine-Grained Sports Action Recognition

ICCV 2021poster

Human pose is a useful feature for fine-grained sports action understanding. However, pose estimators are often unreliable when run on sports video due to domain shift and factors such as motion blur and occlusions. This leads to poor accuracy when downstream tasks, such as action recognition, depen…

Cited by 58PDFcodeScholar