← Search

Qian HE

39 accepted papers

2026

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

ICML 2026poster

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and audio-driven video animation (RA2V) as isolated objectives. F…

Cited by 0SourceScholar
2026

DreamStyle: A Unified Framework for Video Stylization

CVPR 2026

Video stylization, an important downstream task of video generation models, has not yet been thoroughly explored. Its input style conditions typically include text, style image, and stylized first frame. Each condition has a characteristic advantage: text is more flexible, style image provides a mor

Cited by 0SourcecodeScholar
2026

DuRP: Dual-Stage Physics-Embedded Learning for Joint Radiance and Polarization Restoration

ICML 2026poster

Polarization information is valuable for many computer vision applications. However, in hazy environments, polarization information is severely attenuated due to the degradation of captured polarized images. Existing dehazing methods struggle to effectively restore polarization information, as singl…

Cited by 0SourceScholar
2026

FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy

ICML 2026poster

Visuomotor policies aim to learn complex manipulation tasks from expert demonstrations. However, generating smooth and coherent trajectories remains challenging, as it requires balancing proximal precision with distal foresight. Existing approaches typically focus on optimizing intra-chunk action di…

Cited by 0SourceScholar
2026

Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

AAAI 2026technical

Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, images, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of modality-complete data and the difficulty

Cited by 0SourcePDFScholar
2026

Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset

ICLR 2026poster

Subject-to-video generation has witnessed substantial progress in recent years. However, existing models still face significant challenges in faithfully following textual instructions. This limitation, commonly known as the copy-paste problem, arises from the widely used in-pair training paradigm. T…

Cited by 0SourcecodeScholar
2026

Scaling Multi-Identity Consistency for Image Customization via Multi-to-Multi Matching Paradigm

CVPR 2026

Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving consistent identity while avoiding identity confusion with multi-re

Cited by 0SourcecodeScholar
2026

Scaling4D: Pushing the Frontier of Video Novel View Synthesis through Large-Scale Monocular Videos

CVPR 2026

Video Novel View Synthesis (VNVS) aims to render arbitrary novel viewpoints of dynamic scenes from a single-view video, but its algorithmic training faces a major challenge: the lack of large-scale multi-view video datasets. Prior methods often train on monocular data by framing it as an inpainting

Cited by 0SourceScholar
2026

Unified Customized Generation by Disentangled Reward Modeling

CVPR 2026

Existing literature typically treats various customized generation tasks (e.g., subject-customized generation, style-customized generation) as distinct and disjoint problems, with each task focusing solely on customizing a specific aspect of the reference image. However, we argue that the objectives

Cited by 0SourcecodeScholar
2025

AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion

CVPR 2025poster

The task of video generation requires synthesizing visually realistic and temporally coherent video frames. Existing methods primarily use asynchronous auto-regressive models or synchronous diffusion models to address this challenge. However, asynchronous auto-regressive models often suffer from inc…

2025

AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models

CVPR 2025poster

Recent advances in garment-centric image generation from text and image prompts based on diffusion models are impressive. However, existing methods lack support for various combinations of attire, and struggle to preserve the garment details while maintaining faithfulness to the text prompts, limiti…

Cited by 5SourcePDFScholar
2025

GLAM: Global-Local Variation Awareness in Mamba-based World Model

AAAI 2025technical

Mimicking the real interaction trajectory in the inference of the world model has been shown to improve the sample efficiency of model-based reinforcement learning (MBRL) algorithms. Many methods directly use known state sequences for reasoning. However, this approach fails to enhance the quality of…

2025

HyperLoRA: Parameter-Efficient Adaptive Generation for Portrait Synthesis

CVPR 2025highlight

Personalized portrait synthesis, essential in domains like social entertainment, has recently made significant progress. Person-wise fine-tuning based methods, such as LoRA and DreamBooth, can produce photorealistic outputs but need training on individual samples, consuming time and resources and po…

Cited by 0SourcePDFScholar
2025

I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength

ICLR 2025poster

Video generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However, existing camera control methods still suffer from several limita…

2025

I2VControl: Disentangled and Unified Video Motion Synthesis Control

ICCV 2025poster

Motion controllability is crucial in video synthesis. However, most previous methods are limited to single control types, and combining them often results in logical conflicts. In this paper, we propose a disentangled and unified framework, namely I2VControl, to overcome the logical conflicts. We re…

2025

Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

ICCV 2025poster

Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data scalability and subject expansibility. For the first challenge, moving from curating single-subject datasets to multiple-subject ones and scaling them is…

2025

Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

CVPR 2025poster

Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask^2DiT,…

2025

OneGT: One-Shot Geometry-Texture Neural Rendering for Head Avatars

ICCV 2025poster

Existing solutions for creating high-fidelity digital head avatars encounter various obstacles. Traditional rendering tools offer realistic results, while heavily requiring expert skills. Neural rendering methods are more efficient but often compromise between the generated fidelity and flexibility.…

Cited by 0SourcePDFScholar
2025

Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment

ICCV 2025poster

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts subject elements from reference images and generates subject-con…

Cited by 0SourcePDFScholar
2024

DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations

CVPR 2024highlight

The diffusion-based text-to-image model harbors immense potential in transferring reference style. However current encoder-based approaches significantly impair the text controllability of text-to-image models while transferring styles. In this paper we introduce DEADiff to address this issue using…

2024

DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image Generation

AAAI 2024technical

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity and follow the text prompts simultaneously for conditioned input face images and texts. Despite existing encoder-based methods…

Cited by 34SourcePDFScholar
2024

FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion Models

NeurIPS 2024poster

Modeling and producing lifelike clothed human images has attracted researchers' attention from different areas for decades, with the complexity from highly articulated and structured content. Rendering algorithms decompose and simulate the imaging process of a camera, while are limited by the accura…

Cited by 1SourcePDFScholar
2024

PuLID: Pure and Lightning ID Customization via Contrastive Alignment

NeurIPS 2024poster

We propose Pure and Lightning ID customization (PuLID), a novel tuning-free ID customization method for text-to-image generation. By incorporating a Lightning T2I branch with a standard diffusion one, PuLID introduces both contrastive alignment loss and accurate ID loss, minimizing disruption to the…

2024

RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization

CVPR 2024poster

Text-to-image customization which aims to synthesize text-driven images for the given subjects has recently revolutionized content creation. Existing works follow the pseudo-word paradigm i.e. represent the given subjects as pseudo-words and then compose them with the given text. However the inheren…

2023

ReGANIE: Rectifying GAN Inversion Errors for Accurate Real Image Editing

AAAI 2023technical

The StyleGAN family succeed in high-fidelity image generation and allow for flexible and plausible editing of generated images by manipulating the semantic-rich latent style space. However, projecting a real image into its latent space encounters an inherent trade-off between inversion quality and e…

Cited by 7SourcePDFScholar
2023

Semantic 3D-Aware Portrait Synthesis and Manipulation Based on Compositional Neural Radiance Field

AAAI 2023technical

Recently 3D-aware GAN methods with neural radiance field have developed rapidly. However, current methods model the whole image as an overall neural radiance field, which limits the partial semantic editability of synthetic results. Since NeRF renders an image pixel by pixel, it is possible to split…

2023

Target Velocity Estimation for Quantization-Based Cooperative MIMO Radar and Communications System

ICASSP 2023accepted

Target velocity estimation is investigated for a cooperative multiple-input multiple-output (MIMO) integrated radar and communications (IRC) system employing quantized measurements. To reduce the communications burden, the local receivers quantize the local measurements, and then transmit the quanti…

Cited by 0SourceScholar
2023

UGC: Unified GAN Compression for Efficient Image-to-Image Translation

ICCV 2023poster

Recent years have witnessed the prevailing progress of Generative Adversarial Networks (GANs) in image-to-image translation. However, the success of these GAN models hinges on ponderous computational costs and labor-expensive training data. Current efficient GAN learning techniques often fall into t…

Cited by 4PDFcodeScholar
2022

XMP-Font: Self-Supervised Cross-Modality Pre-Training for Few-Shot Font Generation

CVPR 2022poster

Generating a new font library is a very labor-intensive and time-consuming job for glyph-rich scripts. Few-shot font generation is thus required, as it requires only a few glyph references without fine-tuning during test. Existing methods follow the style-content disentanglement paradigm, and expect…

Cited by 58PDFScholar
2021

Parameter Estimation for Coherent Passive MIMO Radar with Unknown Signals under Direct Path Influence

ICASSP 2021accepted

When the radar antennas are properly placed so that each antenna falls within the same target beamwidth, the coherent processing can be employed. This paper studies the problem of joint target position and velocity estimation for a coherent passive radar system. The received observation model with d…

Cited by 0SourceScholar
2020

Distribution of the Product of a Complex Gaussian Matrix and Vector and Its Sum with a Complex Gaussian Vector

ICASSP 2020accepted

In this paper, we derive the distribution of the product of a complex Gaussian matrix and a complex Gaussian vector. Further, we calculate the distribution of the sum of this product and a complex Gaussian vector, which generalizes the recent results where a complex Gaussian scalar is considered ins…

Cited by 0SourceScholar
2019

Target Localization and Mutual Information Improvement for Cooperative MIMO Radar and MIMO Communication Systems

ICASSP 2019accepted

In this work, we study coexisting MIMO radar and MIMO communication systems, where the two systems work cooperatively. The radar shares its antenna positions and transmitted signals with the communication system. The communication system informs the radar about the antenna locations, as well as the…

Cited by 0SourceScholar