Gaussian-Face: Talking Head Generation with Hybrid Density via 3D Gaussian Splatting
Guanwen Feng, Yilin Zhang, Yunan Li, Siyu Jin, Qiguang Miao
Abstract
In recent years, audio-driven neural radiance field (NeRF)-based talking head generation techniques have achieved impressive results. However, these methods still have some limitations, such as unsynchronized lip movements and visual jitter. Recently, 3D Gaussian splatting has gradually replaced NeRF. Compared to NeRF, 3D Gaussian offers notable advantages, including higher efficiency and better reconstruction quality. Based on this, we propose Gaussian-Face, an audio-driven Gaussian-based facial avatar. With just a few minutes of monocular video and audio, a high-fidelity, driveable facial avatar can be reconstructed within hours. To achieve this, we first use FLAME to obtain the 3D representation of the face and then design a Lip Motion Translator to map audio to 3D lip representations. To model higher-quality facial details, we propose a hybrid density modeling method that balances rendering speed and quality, enabling our approach to render high-fidelity facial avatars at more than 160 FPS. Project page: https://peterfanfan.github.io/Gaussian-Face/
BibTeX
@inproceedings{icassp2025_gaussianfacetalk,
title = {Gaussian-Face: Talking Head Generation with Hybrid Density via 3D Gaussian Splatting},
author = {Guanwen Feng and Yilin Zhang and Yunan Li and Siyu Jin and Qiguang Miao},
booktitle = {ICASSP 2025},
year = {2025}
}