← Search

Kaihui Cheng

7 accepted papers

2026

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture

CVPR 2026

This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at the training timestep is the corresponding ground-truth noisy data that is an interpolation of the noise and the data, and duri

Cited by 0SourcecodeScholar
2026

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

ICML 2026poster

Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visual latents throughout denoising. In this setting, we observe a prompt forgetting phenomenon: the semantics of the prompt …

Cited by 0SourceScholar
2025

4D Diffusion for Dynamic Protein Structure Prediction with Reference and Motion Guidance

AAAI 2025technical

Protein structure prediction is pivotal for understanding the structure-function relationship of proteins, advancing biological research, and facilitating pharmaceutical development and experimental design. While deep learning methods and the expanded availability of experimental 3D protein structur…

Cited by 0SourcePDFScholar
2025

Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

ICLR 2025poster

Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities.First, we extend…

2025

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

CVPR 2025poster

Existing methodologies for animating portrait images face significant challenges, particularly in handling non-frontal perspectives, rendering dynamic objects around the portrait, and generating immersive, realistic backgrounds. In this paper, we introduce the first application of a pretrained trans…

2025

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

CVPR 2025highlight

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality, human-centric video datasets presents a challenge to progress in…

Cited by 2SourcePDFScholar
2023

TeAw: Text-Aware Few-Shot Remote Sensing Image Scene Classification

ICASSP 2023accepted

The recent advance has shown that few-shot learning may be a promising way to alleviate the data reliance of remote sensing image scene classification. However, most existing works focus on extracting distinguishable features only from visual modality, while the problem of learning knowledge from mu…

Cited by 0SourceScholar