2024
SYNTHE-SEES: Face Based Text-to-Speech for Virtual Speaker
ICASSP 2024accepted
Recent virtual voice generation researches have limitations in that they results in low-quality voice and generate inconsistent voice from the same speaker’s different facial images. To handle this, we propose a facial encoder module for the pre-trained multi-speaker TTS system called SYNTHE-SEES, w…