Shabdh: A multi lingual zero-shot voice cloning approach with speaker disentanglement
Sreeram Manghat, Sreeja Manghat, Tanja Schultz
Abstract
This paper presents a zero-shot voice cloning system leveraging the DIS-Vector framework, which disentangles and encodes key speech features: content, pitch, timbre, and rhythm. Using the YourTTS architecture, the system synthesizes high-quality speech with precise control over both speaker identity and speech characteristics. The approach integrates multilingual data from the LIMMITS-25 dataset. The system employs neural codec TTS and clustering techniques for efficient and personalized speech synthesis. By utilizing DIS-Vector embeddings, the system enables zero-shot voice cloning, allowing the synthesis of speech in the voice of any unseen speaker with high fidelity and adaptability across multiple languages.
BibTeX
@inproceedings{icassp2025_shabdhamultiling,
title = {Shabdh: A multi lingual zero-shot voice cloning approach with speaker disentanglement},
author = {Sreeram Manghat and Sreeja Manghat and Tanja Schultz},
booktitle = {ICASSP 2025},
year = {2025}
}