← Search

Yabo Zhang

8 accepted papers

2026

PreferThinker: Reasoning-based Personalized Image Preference Assessment

ICLR 2026poster

Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing methods mainly focus on general preference assessment, training models with large-scale data to tackle well-defined task…

Cited by 0SourceScholar
2025

FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors

ICCV 2025poster

Interactive image editing allows users to modify images through visual interaction operations such as drawing, clicking, and dragging. Existing methods construct such supervision signals from videos, as they capture how objects change with various physical interactions. However, these models are usu…

2025

MC^2: Multi-concept Guidance for Customized Multi-concept Generation

CVPR 2025poster

Customized text-to-image generation, which synthesizes images based on user-specified concepts, has made significant progress in handling individual concepts. However, when extended to multiple concepts, existing methods often struggle with properly integrating different models and avoiding the unin…

2025

VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

AAAI 2025technical

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment, owing to insufficient quality and quantity of training videos.…

2024

ControlVideo: Training-free Controllable Text-to-video Generation

ICLR 2024poster

Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart lags behind due to the excessive training cost. To avert the training burden, we propose a training-free ControlVideo to produce high-quality videos based on the provided text prom…

2024

VQ-FONT: Few-Shot Font Generation with Structure-Aware Enhancement and Quantization

AAAI 2024technical

Few-shot font generation is challenging, as it needs to capture the fine-grained stroke styles from a limited set of reference glyphs, and then transfer to other characters, which are expected to have similar styles. However, due to the diversity and complexity of Chinese font styles, the synthesize…

2023

ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

ICCV 2023oral

In addition to the unprecedented ability in imaginary creation, large text-to-image models are expected to take customized concepts in image generation. Existing works generally learn such concepts in an optimization-based manner, yet bringing excessive computation or memory burden. In this paper, w…

Cited by 361PDFcodeScholar
2022

Towards Diverse and Faithful One-shot Adaption of Generative Adversarial Networks

NeurIPS 2022accept

One-shot generative domain adaption aims to transfer a pre-trained generator on one domain to a new domain using one reference image only. However, it remains very challenging for the adapted generator (i) to generate diverse images inherited from the pre-trained generator while (ii) faithfully acqu…