ICASSP 2025accepted0 citations

KAN-Face: Efficient Resource Usage and Precision Lip-Sync in Talking Head Generation

Guanwen Feng, Siyu Jin, Zhihao Qian, Yunan Li, Qiguang Miao

Abstract

Despite significant progress in NeRF-based talking head generation, problems like poor lip synchronization and inefficient resource usage remain. To solve these, we propose KANFace, a lightweight framework. In preprocessing, we introduce a Lip-Sync Enhancement Module that uses Wav2Lip to extract high-resolution audio features and map them to an explicit intermediate representation, ensuring precise lip movement alignment with the speaker’s identity. These predicted lip-sync features are combined with fundamental audio-extracted lip features and injected into the rendering module to improve synchronization. For rendering, we introduce FastKAN to map spatial points to color values. As a variant of KAN, FastKAN’s sensitivity to 3D scenes and efficient structure enable precise, fast color prediction. Our framework reduces resource consumption while enhancing lip-sync accuracy and facial reconstruction, making it ideal for talking head generation tasks in resource-limited settings. Project: https://peterfanfan.github.io/KAN-Face/

BibTeX
@inproceedings{icassp2025_kanfaceefficient,
  title = {KAN-Face: Efficient Resource Usage and Precision Lip-Sync in Talking Head Generation},
  author = {Guanwen Feng and Siyu Jin and Zhihao Qian and Yunan Li and Qiguang Miao},
  booktitle = {ICASSP 2025},
  year = {2025}
}
KAN-Face: Efficient Resource Usage and Precision Lip-Sync in Talking Head Generation · ICASSP 2025