← Search

Anh Tran

24 accepted papers

2026

Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation

CVPR 2026

Despite recent advances in personalized image generation, existing models consistently fail to produce reliable multi-human scenes, often merging or losing facial identity. We present Ar2Can, a novel two-stage framework that disentangles spatial planning from identity rendering for multi-human gener

Cited by 0SourcecodeScholar
2026

InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting

CVPR 2026

Recent diffusion-based models achieve photorealism in image inpainting but require many sampling steps, limiting practical use. Few-step text-to-image models offer faster generation, but naively applying them to inpainting yields poor harmonization and artifacts between the background and inpainted

Cited by 0SourceScholar
2026

PixelRush: Ultra-Fast, Training-Free High-Resolution Image Generation via One-step Diffusion

CVPR 2026

Pre-trained diffusion models excel at generating high-quality images but remain inherently limited by their native training resolution. Recent training-free approaches have attempted to overcome this constraint by introducing interventions during the denoising process; however, these methods incur s

Cited by 0SourceScholar
2025

Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking

CVPR 2025poster

Existing 3D instance segmentation methods frequently encounter issues with over-segmentation, leading to redundant and inaccurate 3D proposals that complicate downstream tasks. This challenge arises from their unsupervised merging approach, where dense 2D instance masks are lifted across frames into…

Cited by 0SourcePDFScholar
2025

CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models

ICCV 2025poster

Disentangling content and style from a single image, known as content-style decomposition (CSD), enables recontextualization of extracted content and stylization of extracted styles, offering greater creative flexibility in visual synthesis. While recent personalization methods have explored the dec…

Cited by 0SourcePDFScholar
2025

Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Image Generation

AAAI 2025technical

Flow matching has emerged as a promising framework for training generative models, demonstrating impressive empirical performance while offering relative ease of training compared to diffusion-based models. However, this method still requires numerous function evaluations in the sampling process. To…

2025

SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models

ICCV 2025poster

The rapid growth of text-to-image diffusion models has raised concerns about their potential misuse in generat- ing harmful or unauthorized contents. To address these issues, several Concept Erasure methods have been pro- posed. However, most of them fail to achieve both robust- ness, i.e., the abil…

Cited by 0SourcePDFScholar
2025

Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts

ICCV 2025poster

The escalating demand for real-time image synthesis has driven significant advancements in one-step diffusion models, which inherently offer expedited generation speeds compared to traditional multi-step methods. However, this enhanced efficiency is frequently accompanied by a compromise in the cont…

Cited by 0SourcePDFScholar
2025

SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion

CVPR 2025poster

Recent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-based text-to-image models. However, these methods often fall short of the speed demands required for real-world and on-device applicat…

2024

Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown Domains

CVPR 2024poster

This paper presents an innovative framework designed to train an image deblurring algorithm tailored to a specific camera device. This algorithm works by transforming a blurry input image which is challenging to deblur into another blurry image that is more amenable to deblurring. The transformation…

2024

COMBAT: Alternated Training for Effective Clean-Label Backdoor Attacks

AAAI 2024technical

Backdoor attacks pose a critical concern to the practice of using third-party data for AI development. The data can be poisoned to make a trained model misbehave when a predefined trigger pattern appears, granting the attackers illegal benefits. While most proposed backdoor attacks are dirty-label,…

2024

Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance

CVPR 2024poster

We introduce Open3DIS a novel solution designed to tackle the problem of Open-Vocabulary Instance Segmentation within 3D scenes. Objects within 3D environments exhibit diverse shapes scales and colors making precise instance-level identification a challenging task. Recent advancements in Open-Vocabu…

2024

SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation

CVPR 2024poster

Despite their ability to generate high-resolution and diverse images from text prompts text-to-image diffusion models often suffer from slow iterative sampling processes. Model distillation is one of the most effective directions to accelerate these models. However previous distillation methods fail…

2023

Anti-DreamBooth: Protecting Users from Personalized Text-to-image Synthesis

ICCV 2023poster

Text-to-image diffusion models are nothing but a revolution, allowing anyone, even without design skills, to create realistic images from simple text inputs. With powerful personalization tools like DreamBooth, they can generate images of a specific person just by learning from his/her few reference…

Cited by 127PDFcodeScholar
2023

Efficient Scale-Invariant Generator With Column-Row Entangled Pixel Synthesis

CVPR 2023poster

Any-scale image synthesis offers an efficient and scalable solution to synthesize photo-realistic images at any scale, even going beyond 2K resolution. However, existing GAN-based solutions depend excessively on convolutions and a hierarchical architecture, which introduce inconsistency and the "tex…

2023

HyperCUT: Video Sequence From a Single Blurry Image Using Unsupervised Ordering

CVPR 2023poster

We consider the challenging task of training models for image-to-video deblurring, which aims to recover a sequence of sharp images corresponding to a given blurry image input. A critical issue disturbing the training of an image-to-video model is the ambiguity of the frame ordering since both the f…