ICASSP 2025accepted0 citations

Fine-portraitist: Visualizing the Speaker's Face Portrait during Speech Listening

Jinting Wang, Li Liu, Jun Wang

Abstract

Speech-to-portrait generation (S2P) plays a crucial role in speech-driven, human-centered creative content generation, aiming to synthesize a speaker’s face portrait with identity consistency from a given speech clip. However, existing S2P methods can typically only preserve attribute consistency, e.g., gender and age, while failing to capture the more important part-appearance consistency due to the coarse speech-face correlation. In this work, we propose Fine-portraitist, a novel retrieval-augmented, easy-to-hard generation framework designed to tackle this problem. Specifically, Fine-portraitist enhances identity consistency in S2P through two key innovations: 1) We first explore the fine-grained speech-face correlation by decomposing the face portrait into speech-related and speech-unrelated parts. Based on this, we propose a two-stage, diffusion-based pipeline to progressively achieve S2P; 2) A retrieval prior is introduced, selected from a retrieval database based on speech feature similarity, providing supplementary external information for more accurate and realistic generation results. Extensive experiments on two datasets, i.e., AVSpeech and VoxCeleb, demonstrate that Fine-portraitist significantly outperforms existing S2P methods.

BibTeX
@inproceedings{icassp2025_fineportraitistv,
  title = {Fine-portraitist: Visualizing the Speaker's Face Portrait during Speech Listening},
  author = {Jinting Wang and Li Liu and Jun Wang},
  booktitle = {ICASSP 2025},
  year = {2025}
}