Lightweight Prosody-TTS for Multi-Lingual Multi-Speaker Scenario
Giridhar Pamisetty, Sahukari Chaitanya Varun, K. Sri Rama Murty
Abstract
This work presents a lightweight end-to-end text-to-speech (TTS) synthesis for the multi-lingual multi-speaker (ML-MS) scenario. The proposed system uses nonautoregressive modular architecture with interconnected subnets for text-encoder, duration estimator, f<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> estimator, and acoustic decoder. The text encoder is conditioned with language embeddings, while the duration and f<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> estimators are conditioned with speaker embeddings. All the subnets are optimized in an end-to-end fashion using accumulated loss across the modules. The intermediate auxiliary loss functions help effectively capture the speech information with lesser data. The proposed model achieved a mean opinion score (MOS) of 4.40 and a speaker similarity score of 3.8 with just 4.89 million (M) parameters in LIMMITS grand challenge organized as part of ICASSP-23.
BibTeX
@inproceedings{icassp2023_lightweightproso,
title = {Lightweight Prosody-TTS for Multi-Lingual Multi-Speaker Scenario},
author = {Giridhar Pamisetty and Sahukari Chaitanya Varun and K. Sri Rama Murty},
booktitle = {ICASSP 2023},
year = {2023}
}