← Search

Chaoyue Wang

23 accepted papers

2025

MagicNaming: Consistent Identity Generation by Finding a “Name Space” in T2I Diffusion Models

AAAI 2025technical

Large-scale text-to-image diffusion models, (e.g., DALL-E, SDXL) are capable of generating famous persons by simply referring to their names. Is it possible to make such models generate generic identities as simple as the famous ones, e.g., just use a name? In this paper, we explore the existence of…

Cited by 1SourcePDFScholar
2025

Semantix: An Energy-guided Sampler for Semantic Style Transfer

ICLR 2025poster

Recent advances in style and appearance transfer are impressive, but most methods isolate global style and local appearance transfer, neglecting semantic correspondence. Additionally, image and video tasks are typically handled in isolation, with little focus on integrating them for video transfer.…

Cited by 0SourcePDFScholar
2024

Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation

AAAI 2024technical

A 360-degree (omni-directional) image provides an all-encompassing spherical view of a scene. Recently, there has been an increasing interest in synthesising 360-degree images from conventional narrow field of view (NFoV) images captured by digital cameras and smartphones, for providing immersive ex…

2024

Decomposing Semantic Shifts for Composed Image Retrieval

AAAI 2024technical

Composed image retrieval is a type of image retrieval task where the user provides a reference image as a starting point and specifies a text on how to shift from the starting point to the desired target image. However, most existing methods focus on the composition learning of text and reference im…

2024

Dual Mapping of 2D StyleGAN for 3D-Aware Image Generation and Manipulation (Student Abstract)

AAAI 2024technical

3D-aware GANs successfully solve the problem of 3D-consistency generation and furthermore provide a 3D shape of the generated object. However, the application of the volume renderer disturbs the disentanglement of the latent space, which makes it difficult to manipulate 3D-aware GANs and lowers the…

Cited by 0SourcePDFScholar
2024

Eliminating the Cross-Domain Misalignment in Text-guided Image Inpainting

IJCAI 2024poster

Text-guided image inpainting has rapidly garnered prominence as a task in user-directed image synthesis, aiming to complete the occluded image regions following the textual prompt provided. However, current methods usually grapple with issues arising from the disparity between low-level pixel data a…

2024

Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion Models

AAAI 2024technical

Denoising Diffusion Probabilistic Models (DDPMs) have achieved significant success in generation tasks. Nevertheless, the exposure bias issue, i.e., the natural discrepancy between the training (the output of each step is calculated individually by a given input) and inference (the output of each st…

Cited by 2SourcePDFScholar
2024

One More Step: A Versatile Plug-and-Play Module for Rectifying Diffusion Schedule Flaws and Enhancing Low-Frequency Controls

CVPR 2024poster

It is well known that many open-released foundational diffusion models have difficulty in generating images that substantially depart from average brightness despite such images being present in the training data. This is due to an inconsistency: while denoising starts from pure Gaussian noise durin…

Cited by 3SourcePDFScholar
2023

All Points Matter: Entropy-Regularized Distribution Alignment for Weakly-supervised 3D Segmentation

NeurIPS 2023poster

Pseudo-labels are widely employed in weakly supervised 3D segmentation tasks where only sparse ground-truth labels are available for learning. Existing methods often rely on empirical label selection strategies, such as confidence thresholding, to generate beneficial pseudo-labels for model training…

2023

Cocktail: Mixing Multi-Modality Control for Text-Conditional Image Generation

NeurIPS 2023poster

Text-conditional diffusion models are able to generate high-fidelity images with diverse contents. However, linguistic representations frequently exhibit ambiguous descriptions of the envisioned objective imagery, requiring the incorporation of additional control signals to bolster the efficacy of t…

Cited by 23SourcePDFScholar
2023

Domain Re-Modulation for Few-Shot Generative Domain Adaptation

NeurIPS 2023poster

In this study, we delve into the task of few-shot Generative Domain Adaptation (GDA), which involves transferring a pre-trained generator from one domain to a new domain using only a few reference images. Inspired by the way human brains acquire knowledge in new domains, we present an innovative gen…

2023

MagicFusion: Boosting Text-to-Image Generation Performance by Fusing Diffusion Models

ICCV 2023poster

The advent of open-source AI communities has produced a cornucopia of powerful text-guided diffusion models that are trained on various datasets. While few explorations have been conducted on ensembling such models to combine their strengths. In this work, we propose a simple yet effective method ca…

Cited by 16PDFcodeScholar
2023

Unified Discrete Diffusion for Simultaneous Vision-Language Generation

ICLR 2023poster

The recently developed discrete diffusion model performs extraordinarily well in generation tasks, especially in the text-to-image task, showing great potential for modeling multimodal signals. In this paper, we leverage these properties and present a unified multimodal generation model, which can p…

2022

FakeCLR: Exploring Contrastive Learning for Solving Latent Discontinuity in Data-Efficient GANs

ECCV 2022poster

"Data-Efficient GANs (DE-GANs), which aim to learn generative models with a limited amount of training data, encounter several challenges for generating high-quality samples. Since data augmentation strategies have largely alleviated the training instability, how to further improve the generative pe…

2022

Modeling Image Composition for Complex Scene Generation

CVPR 2022poster

We present a method that achieves state-of-the-art results on challenging (few-shot) layout-to-image generation tasks by accurately modeling textures, structures and relationships contained in a complex scene. After compressing RGB images into patch tokens, we propose the Transformer with Focal Atte…

Cited by 57PDFcodeScholar
2022

Self-Augmented Unpaired Image Dehazing via Density and Depth Decomposition

CVPR 2022poster

To overcome the overfitting issue of dehazing models trained on synthetic hazy-clean image pairs, many recent methods attempted to improve models' generalization ability by training on unpaired data. Most of them simply formulate dehazing and rehazing cycles, yet ignore the physical properties of th…

Cited by 261PDFcodeScholar
2022

SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

NeurIPS 2022accept

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (MAE) different between vision and language. In this paper, we explore a potential…

2022

Visual Semantics Allow for Textual Reasoning Better in Scene Text Recognition

AAAI 2022technical

Existing Scene Text Recognition (STR) methods typically use a language model to optimize the joint probability of the 1D character sequence predicted by a visual recognition (VR) model, which ignore the 2D spatial context of visual semantics within and between character instances, making them not ge…

2020

FeatureFlow: Robust Video Interpolation via Structure-to-Texture Generation

CVPR 2020poster

Video interpolation aims to synthesize non-existent frames between two consecutive frames. Although existing optical flow based methods have achieved promising results, they still face great challenges in dealing with the interpolation of complicated dynamic scenes, which include occlusion, blur or…

Cited by 96PDFcodeScholar
2020

PuppeteerGAN: Arbitrary Portrait Animation With Semantic-Aware Appearance Transformation

CVPR 2020poster

Portrait animation, which aims to animate a still portrait to life using poses extracted from target frames, is an important technique for many real-world entertainment applications. Although recent works have achieved highly realistic results on synthesizing or controlling human head images, the pu…

Cited by 58PDFScholar