← Search

Shaozhe Hao

7 accepted papers

2025

ArtiFade: Learning to Generate High-quality Subject from Blemished Images

CVPR 2025poster

Subject-driven text-to-image generation has demonstrated remarkable advancements in its ability to learn and capture characteristics of a subject using only a limited number of images. However, existing methods commonly rely on high-quality images for training and often struggle to generate reasonab…

Cited by 1SourcePDFScholar
2025

BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities

ICLR 2025poster

We introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation capabilities. BiGR is the first conditional generative model that unifies generation and discrimination within the same fr…

2025

Elucidating the design space of language models for image generation

ICML 2025poster

The success of large language models (LLMs) in text generation has inspired their application to image generation. However, existing methods either rely on specialized designs with inductive biases or adopt LLMs without fully exploring their potential in vision tasks. In this work, we systematically…

2025

Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

NeurIPS 2025poster

Video content editing has a wide range of applications. With the advancement of diffusion-based generative models, video editing techniques have made remarkable progress, yet they still remain far from practical usability. Existing inversion-based video editing methods are time-consuming and struggl…

Cited by 0SourcecodeScholar
2024

Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation

ECCV 2024poster

"Text-to-image generation has made significant advancements with the introduction of text-to-image diffusion models. These models typically consist of a language model that interprets user prompts and a vision model that generates corresponding images. As language and vision models continue to progr…

2023

Learning Attention As Disentangler for Compositional Zero-Shot Learning

CVPR 2023poster

Compositional zero-shot learning (CZSL) aims at learning visual concepts (i.e., attributes and objects) from seen compositions and combining concept knowledge into unseen compositions. The key to CZSL is learning the disentanglement of the attribute-object composition. To this end, we propose to exp…

2023

Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models

NeurIPS 2023poster

Text-to-Image diffusion models have made tremendous progress over the past two years, enabling the generation of highly realistic images based on open-domain text descriptions. However, despite their success, text descriptions often struggle to adequately convey detailed controls, even when composed…