← Search

Quan Dao

8 accepted papers

2026

An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning

CVPR 2026

In online incremental learning, data continuously arrives with substantial shifts in distribution, creating a significant challenge since previous samples cannot be revisited. Prior research has typically relied on either a single adaptive centroid or fixed multiple centroids to represent each class

Cited by 0SourceScholar
2026

Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model

CVPR 2026

Transformer architectures, particularly Diffusion Transformers (DiTs), have become widely used in diffusion and flow-matching models due to their strong performance compared to convolutional UNets. However, the isotropic design of DiTs processes the same number of patchified tokens in every block, l

Cited by 0SourcecodeScholar
2025

AutoEdit: Automatic Hyperparameter Tuning for Image Editing

NeurIPS 2025poster

Recent advances in diffusion models have revolutionized text-guided image editing, yet existing editing methods face critical challenges in hyperparameter identification. To get the reasonable editing performance, these methods often require the user to brute-force tune multiple interdependent hyper…

Cited by 0SourceScholar
2025

Improved Training Technique for Latent Consistency Models

ICLR 2025poster

Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demonstrated impressive performance, achieving results on par with diffusion models in the pixel space. However, the success…

2025

Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Image Generation

AAAI 2025technical

Flow matching has emerged as a promising framework for training generative models, demonstrating impressive empirical performance while offering relative ease of training compared to diffusion-based models. However, this method still requires numerous function evaluations in the sampling process. To…

2024

DiMSUM: Diffusion Mamba - A Scalable and Unified Spatial-Frequency Method for Image Generation

NeurIPS 2024poster

We introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in input images for image generation tasks. While state-space networks, including Mamba, a revolutionary advancement in re…

2023

Anti-DreamBooth: Protecting Users from Personalized Text-to-image Synthesis

ICCV 2023poster

Text-to-image diffusion models are nothing but a revolution, allowing anyone, even without design skills, to create realistic images from simple text inputs. With powerful personalization tools like DreamBooth, they can generate images of a specific person just by learning from his/her few reference…

Cited by 127PDFcodeScholar