Diffused Poses and Distilled Expressions for Controllable Audio-driven Talking Face Generation
Ziqi Zhou, Weize Quan, Zhaojin Lu, Dong-Ming Yan
Abstract
Audio-driven portrait animation is an emerging field in multi-modal generation that aims to create lifelike talking face videos from audio input. While significant progress has been made, accurately modeling the relationship between audio signals and various facial motions, such as head poses and expressions, remains a challenge. Existing methods have primarily focused on generating lip-synchronized movements, often neglecting the intricate correlations between audio and other facial dynamics like head movements and eye blinks. More recent approaches have attempted to address these limitations by introducing latent disentanglement of facial motions, though this often comes at the cost of reduced flexibility in motion control. In this work, we propose a novel framework for audio-driven talking portrait animation that allows for precise and controllable generation of head poses and facial expressions. Our approach includes two key components: an audio-conditional diffusion model for generating prosody-aware head poses and a noise-conditional, lip-distilling transformer for predicting synchronized facial expressions. We further introduce an innovative animation model that uses these generated poses and expressions to produce highly realistic and controllable talking head videos. Extensive experiments demonstrate that our method not only achieves superior performance in generating natural and synchronized facial motions but also outperforms state-of-the-art techniques in the field.
BibTeX
@inproceedings{icassp2025_diffusedposesand,
title = {Diffused Poses and Distilled Expressions for Controllable Audio-driven Talking Face Generation},
author = {Ziqi Zhou and Weize Quan and Zhaojin Lu and Dong-Ming Yan},
booktitle = {ICASSP 2025},
year = {2025}
}