← Search

Ziqiao Peng

12 accepted papers

2026

ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars

CVPR 2026

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We pres

Cited by 0SourceScholar
2026

Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots

RSS 2026poster

Animatronic faces are a central component of socially interactive robots, enabling rich nonverbal communication through facial articulation. However, state-of-the-art animatronic faces are typically tailored systems: each new facial geometry requires extensive manual mechanical redesign, making larg…

Cited by 0SourceScholar
2026

StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars

CVPR 2026

Real-time, streaming interactive avatars represent a critical yet challenging goal in digital human research. Although diffusion-based human avatar generation methods achieve remarkable success, their non-causal architecture and high computational costs make them unsuitable for streaming. Moreover,

Cited by 0SourcecodeScholar
2026

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

CVPR 2026

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a unified framework for human-centric joint audio and video ge

Cited by 0SourceScholar
2025

CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting

IROS 2025

Vehicle-to-everything (V2X) communication plays a crucial role in autonomous driving, enabling cooperation between vehicles and infrastructure. While simulation has significantly contributed to various autonomous driving tasks, its potential for data generation and augmentation in V2X scenarios rema

Cited by 4SourcecodeScholar
2025

DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations

CVPR 2025poster

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward…

2025

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

ICCV 2025poster

Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head rotations and out-of-distribution (OOD) audio. Moreover, they are…

Cited by 0SourcePDFScholar
2025

MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation

NeurIPS 2025poster

Music-driven 3D dance generation has attracted increasing attention in recent years, with promising applications in choreography, virtual reality, and creative content creation. Previous research has generated promising realistic dance movement from audio signals. However, traditional methods underu…

Cited by 0SourceScholar
2025

Morpheus: A Neural-driven Animatronic Face with Hybrid Actuation and Diverse Emotion Control

RSS 2025poster

Previous animatronic faces struggle to effectively express emotions due to both hardware and software limitations. On the hardware side, earlier approaches either used rigid-driven mechanisms, which provide precise control but are difficult to design within constrained spaces, or tendon-driven mecha…

Cited by 0PDFScholar
2025

OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers

NeurIPS 2025spotlight

Lip synchronization is the task of aligning a speaker’s lip movements in video with corresponding speech audio, and it is essential for creating realistic, expressive video content. However, existing methods often rely on reference frames and masked-frame inpainting, which limit their robustness to…

Cited by 0SourceScholar
2024

SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis

CVPR 2024poster

Achieving high synchronization in the synthesis of realistic speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity while Neural Radiance Fields (NeRF) methods although they can address thi…

2023

EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

ICCV 2023poster

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content. To address this issue, this paper proposes an end-to-end neur…

Cited by 117PDFcodeScholar