RQTalker: Speech-driven 3D Facial Animation via Region-aware Vector Quantization
Yingying Fan, Kaisiyuan Wang, Hang Zhou, Shengyi He, Yu Wu
Abstract
Speech-driven 3D facial animation has been a long-standing topic due to the complex geometry and motion modeling as well as difficulties in cross-modality learning. Current studies struggle to synthesize human-like lip motions, as they usually represent the movement of the entire face with a compressed global vector, leading to subtle motion loss and thus over-smoothed movements in the local lip region. To cope with this problem, we propose a new speech-driven 3D facial animation framework RQTalker based on the Region-aware Vector Quantization mechanism. The key insight is to first build a region-aware codebook via a self-reconstruction manner, in which each part of the codebook physically corresponds to a facial region with a clear semantic. Our region-aware codebook divides facial movements into local regions for multiple sub-encodings, reducing information loss from compression and improving local facial motion modeling. In addition, we further propose a spatial-temporal Audio-to-Motion Learning Module to produce movements that are spatially more accurate and temporally consistent. Qualitative and quantitative results demonstrate that our method outperforms state-of-the-art approaches.
BibTeX
@inproceedings{icassp2025_rqtalkerspeechdr,
title = {RQTalker: Speech-driven 3D Facial Animation via Region-aware Vector Quantization},
author = {Yingying Fan and Kaisiyuan Wang and Hang Zhou and Shengyi He and Yu Wu},
booktitle = {ICASSP 2025},
year = {2025}
}