Voice Conversion for Low-Resource Languages via Knowledge Transfer and Domain-Adversarial Training
Huu Tuong Tu, Luong Thanh Long, Vu Huan, Nguyen Thi Phuong Thao, Nguyen Van Thang, Nguyen Tien Cuong, Nguyen Thi Thu Trang
Abstract
Voice conversion (VC) aims to transform speech from a source speaker to a target speaker while preserving the original linguistic content. However, existing VC models typically require large annotated datasets containing transcripts and speaker labels, posing significant challenges for low-resource languages. In any-to-any voice conversion (VC) models that do not rely on annotated datasets, disentangling speaker information from the source speech while preserving linguistic content remains a significant challenge, often resulting in outputs that retain attributes of the source speaker. This paper introduces a novel low-resource VC model that combines knowledge transfer with domain-adversarial training to leverage information from high-resource languages for the benefit of low-resource languages. The proposed approach utilizes pretrained models, with domain-adversarial training enabling the separation of content from speaker identity without the need for annotated datasets. Objective and subjective evaluations on the low-resource Vietnamese language demonstrate that the proposed model outperforms existing methods in terms of naturalness and speaker similarity in low-resource scenarios. Audio samples are available at: https://huutuongtu.github.io/KTVC/index.html.
BibTeX
@inproceedings{icassp2025_voiceconversionf,
title = {Voice Conversion for Low-Resource Languages via Knowledge Transfer and Domain-Adversarial Training},
author = {Huu Tuong Tu and Luong Thanh Long and Vu Huan and Nguyen Thi Phuong Thao and Nguyen Van Thang and Nguyen Tien Cuong and Nguyen Thi Thu Trang},
booktitle = {ICASSP 2025},
year = {2025}
}