← Search

Duomin Wang

9 accepted papers

2026

LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence

ICLR 2026poster

The reliance on implicit point matching via attention has become a core bottleneck in drag-based editing, resulting in a fundamental compromise on weakened inversion strength and costly test-time optimization (TTO). This compromise severely limits the generative capabilities, suppressing high-fidel…

Cited by 0SourceScholar
2026

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

ICLR 2026poster

The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the next major challenge: audio-visual dyadic interactive virtual…

Cited by 0SourcecodeScholar
2026

Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer

ICLR 2026poster

Text-guided color editing in images and videos is a fundamental yet unsolved problem, requiring fine-grained manipulation of color attributes, including albedo, light source color, and ambient lighting, while preserving physical consistency in geometry, material properties, and light-matter interact…

Cited by 0SourceScholar
2025

Taming Teacher Forcing for Masked Autoregressive Video Generation

CVPR 2025poster

We introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation, Complete Teacher Forcing (CTF), conditions masked frames on complete observation frames rather than masked ones (namely…

Cited by 3SourcePDFScholar
2024

PICTURE: PhotorealistIC virtual Try-on from UnconstRained dEsigns

CVPR 2024poster

In this paper we propose a novel virtual try-on from unconstrained designs (ucVTON) task to enable photorealistic synthesis of personalized composite clothing on input human images. Unlike prior arts constrained by specific input types our method allows flexible specification of style (text or image…

Cited by 8SourcePDFScholar
2024

Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer

ECCV 2024poster

"In this paper, we propose a novel learning approach for feed-forward one-shot 4D head avatar synthesis. Different from existing methods that often learn from reconstructing monocular videos guided by 3DMM, we employ pseudo multi-view videos to learn a 4D head synthesizer in a data-driven manner, av…

2024

Portrait4D: Learning One-Shot 4D Head Avatar Synthesis using Synthetic Data

CVPR 2024poster

Existing one-shot 4D head synthesis methods usually learn from monocular videos with the aid of 3DMM reconstruction yet the latter is evenly challenging which restricts them from reasonable 4D head synthesis. We present a method to learn one-shot 4D head synthesis via large-scale synthetic data. The…

Cited by 19SourcePDFScholar
2023

Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head Synthesis

CVPR 2023poster

We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disentangled latent representations and leverage an image generator to synthesize tal…

2023

Talking Head Generation with Probabilistic Audio-to-Visual Diffusion Priors

ICCV 2023poster

We introduce a novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead sample all holistic lip-irrelevant facial motions (i.e. pose, expression, blink, gaze, etc.) to…

Cited by 41PDFScholar