ATP-TTS: Adaptive Thresholding Pseudo-Labeling for Low-Resource Multi-Speaker Text-to-Speech
Feng Li, Shen Chen, Hanjin Yang, Shupei Yuan
Abstract
To address the challenge of high annotation costs in text-to-speech (TTS) generation, this paper introduces a semi-supervised learning framework specifically designed for low-resource TTS scenarios. The framework incorporates adaptive thresholding to select appropriate pseudo-labels and leverages automatic speech recognition (ASR) results, enhanced by contrastive learning perturbations, to predict latent representations. This approach improves the stability and effectiveness of pre-training, allowing for successful transfer learning with only a limited amount of labeled data. Experimental results demonstrate that the proposed method outperforms baseline models in both naturalness and speaker similarity.
BibTeX
@inproceedings{icassp2025_atpttsadaptiveth,
title = {ATP-TTS: Adaptive Thresholding Pseudo-Labeling for Low-Resource Multi-Speaker Text-to-Speech},
author = {Feng Li and Shen Chen and Hanjin Yang and Shupei Yuan},
booktitle = {ICASSP 2025},
year = {2025}
}