ICASSP 2025accepted0 citations

RQTalker: Speech-driven 3D Facial Animation via Region-aware Vector Quantization

Yingying Fan, Kaisiyuan Wang, Hang Zhou, Shengyi He, Yu Wu

Abstract

Speech-driven 3D facial animation has been a long-standing topic due to the complex geometry and motion modeling as well as difficulties in cross-modality learning. Current studies struggle to synthesize human-like lip motions, as they usually represent the movement of the entire face with a compressed global vector, leading to subtle motion loss and thus over-smoothed movements in the local lip region. To cope with this problem, we propose a new speech-driven 3D facial animation framework RQTalker based on the Region-aware Vector Quantization mechanism. The key insight is to first build a region-aware codebook via a self-reconstruction manner, in which each part of the codebook physically corresponds to a facial region with a clear semantic. Our region-aware codebook divides facial movements into local regions for multiple sub-encodings, reducing information loss from compression and improving local facial motion modeling. In addition, we further propose a spatial-temporal Audio-to-Motion Learning Module to produce movements that are spatially more accurate and temporally consistent. Qualitative and quantitative results demonstrate that our method outperforms state-of-the-art approaches.

BibTeX
@inproceedings{icassp2025_rqtalkerspeechdr,
  title = {RQTalker: Speech-driven 3D Facial Animation via Region-aware Vector Quantization},
  author = {Yingying Fan and Kaisiyuan Wang and Hang Zhou and Shengyi He and Yu Wu},
  booktitle = {ICASSP 2025},
  year = {2025}
}