← Search

Weizhi Zhong

4 accepted papers

2026

Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter

ICLR 2026poster

Personalized text-to-image generation aims to synthesize images of user-provided concepts in diverse contexts. Despite recent progress in multi-concept personalization, most are limited to object concepts and struggle to customize abstract concepts (e.g., pose, lighting). Some methods have begun ex…

Cited by 0SourcecodeScholar
2026

SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling

CVPR 2026

Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span different depth layers or are partially occluded. Most image editing methods focus solely on prior information from 2D dataset

Cited by 0SourceScholar
2025

LLM-driven Multimodal and Multi-Identity Listening Head Generation

CVPR 2025poster

Generating natural listener responses in conversational scenarios is crucial for creating engaging digital humans and avatars. Recent work has shown that large language models (LLMs) can be effectively leveraged for this task, demonstrating remarkable capabilities in generating contextually appropri…

Cited by 0SourcePDFScholar
2023

Identity-Preserving Talking Face Generation With Landmark and Appearance Priors

CVPR 2023poster

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have difficulty in generating realistic and lip-synced videos whi…