SYNTHE-SEES: Face Based Text-to-Speech for Virtual Speaker
Jae Hyun Park, Joon-Gyu Maeng, Taejun Bak, Young-Sun Joo
Abstract
Recent virtual voice generation researches have limitations in that they results in low-quality voice and generate inconsistent voice from the same speaker’s different facial images. To handle this, we propose a facial encoder module for the pre-trained multi-speaker TTS system called SYNTHE-SEES, which utilizes face embeddings as speaker embeddings by sharing the embedding space of the pre-trained speech embeddings using cross-modal contrastive learning. We trained the facial encoder in two ways: 1) for consistent embeddings, we use the dataset supervision to capture discriminative speaker attributes; 2) we leverage internal structure of the speech embedding to generate diverse and high-quality voices. Experimental results demonstrate that our method generates more distinct, consistent, and high-quality speaker embeddings than other state-of-the-art methods in both quantitative and qualitative evaluations. Especially, the result of cluster-level evaluation verifies that our method shows the highest distinction performance of diverse speaker embedding. Our demo is available at ${\color{Cyan}{\text{Demo}}}$.
BibTeX
@inproceedings{icassp2024_syntheseesfaceba,
title = {SYNTHE-SEES: Face Based Text-to-Speech for Virtual Speaker},
author = {Jae Hyun Park and Joon-Gyu Maeng and Taejun Bak and Young-Sun Joo},
booktitle = {ICASSP 2024},
year = {2024}
}