Vani: Very-Lightweight Accent-Controllable TTS for Native And Non-Native Speakers With Identity Preservation
Rohan Badlani, Akshit Arora, Subhankar Ghosh, Rafael Valle, Kevin J. Shih, João Felipe Santos, Boris Ginsburg, Bryan Catanzaro
Abstract
We introduce VANI, a very lightweight multi-lingual accent controllable speech synthesis system. Our model builds upon disentanglement strategies proposed in RADMMM[1] and supports explicit control of accent, language, speaker and fine-grained F<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> and energy features for speech synthesis. We utilize the Indic languages dataset, released for LIMMITS 2023 as part of ICASSP Signal Processing Grand Challenge, to synthesize speech in 3 different languages. Our model supports transferring the language of a speaker while retaining their voice and the native accent of the target language. We utilize the large-parameter RADMMM model for Track 1 and lightweight VANI model for Track 2 and 3 of the competition.
BibTeX
@inproceedings{icassp2023_vaniverylightwei,
title = {Vani: Very-Lightweight Accent-Controllable TTS for Native And Non-Native Speakers With Identity Preservation},
author = {Rohan Badlani and Akshit Arora and Subhankar Ghosh and Rafael Valle and Kevin J. Shih and João Felipe Santos and Boris Ginsburg and Bryan Catanzaro},
booktitle = {ICASSP 2023},
year = {2023}
}