ICASSP 2025accepted0 citations

Realistic Real-Time Talking Head Synthesis with Grid Encoding and Progressive Conditioning

Zhiling Ye, Liang-Guo Zhang, Dingheng Zeng, Quan Lu, Ning Jiang

Abstract

Dynamic NeRFs have recently been used for 3D talking portrait synthesis, but challenges remain in improving efficiency and effectiveness. We introduce R2-Talker, an efficient and effective framework for real-time talking head synthesis. Using multi-resolution hash grids, we losslessly encode facial landmarks as conditional features, aligning the structure with facial expression movement. We also propose progressive multilayer conditioning for effective conditional feature fusion. Compared to state-of-the-art works, our approach has superior visual quality and accuracy, and is computationally efficient.

BibTeX
@inproceedings{icassp2025_realisticrealtim,
  title = {Realistic Real-Time Talking Head Synthesis with Grid Encoding and Progressive Conditioning},
  author = {Zhiling Ye and Liang-Guo Zhang and Dingheng Zeng and Quan Lu and Ning Jiang},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Realistic Real-Time Talking Head Synthesis with Grid Encoding and Progressive Conditioning · ICASSP 2025