Fine-portraitist: Visualizing the Speaker's Face Portrait during Speech Listening
Jinting Wang, Li Liu, Jun Wang
Abstract
Speech-to-portrait generation (S2P) plays a crucial role in speech-driven, human-centered creative content generation, aiming to synthesize a speaker’s face portrait with identity consistency from a given speech clip. However, existing S2P methods can typically only preserve attribute consistency, e.g., gender and age, while failing to capture the more important part-appearance consistency due to the coarse speech-face correlation. In this work, we propose Fine-portraitist, a novel retrieval-augmented, easy-to-hard generation framework designed to tackle this problem. Specifically, Fine-portraitist enhances identity consistency in S2P through two key innovations: 1) We first explore the fine-grained speech-face correlation by decomposing the face portrait into speech-related and speech-unrelated parts. Based on this, we propose a two-stage, diffusion-based pipeline to progressively achieve S2P; 2) A retrieval prior is introduced, selected from a retrieval database based on speech feature similarity, providing supplementary external information for more accurate and realistic generation results. Extensive experiments on two datasets, i.e., AVSpeech and VoxCeleb, demonstrate that Fine-portraitist significantly outperforms existing S2P methods.
BibTeX
@inproceedings{icassp2025_fineportraitistv,
title = {Fine-portraitist: Visualizing the Speaker's Face Portrait during Speech Listening},
author = {Jinting Wang and Li Liu and Jun Wang},
booktitle = {ICASSP 2025},
year = {2025}
}