ICASSP 2026poster0 citations

TMD-TTS: A UNIFIED TIBETAN MULTI-DIALECT TEXT-TO-SPEECH FRAMEWORK FOR U-TSANG, AMDO AND KHAM SPEECH DATASET GENERATION

Yutong Liu, Renzeng Duojie, Yuqing Cai, Cheng Huang, Nyima Tashi

Abstract

Tibetan is a low-resource language with limited parallel speech corpora spanning its three major dialects (Ü-Tsang, Amdo, and Kham), limiting progress in speech modeling. To address this issue, we propose TMD-TTS, a unified Tibetan multi-dialect text-to-speech (TTS) framework that synthesizes parallel dialectal speech from explicit dialect labels. Our method features a dialect fusion module and a Dialect-Specialized Dynamic Routing Network (DSDR-Net) to capture fine-grained acoustic and linguistic variations across dialects. Extensive objective and subjective evaluations demonstrate that TMD-TTS significantly outperforms baselines in dialectal expressiveness. We further validate the quality and utility of the synthesized speech through a challenging Speech-to-Speech Dialect Conversion (S2SDC) task.

BibTeX
@inproceedings{icassp2026_tmdttsaunifiedti,
  title = {TMD-TTS: A UNIFIED TIBETAN MULTI-DIALECT TEXT-TO-SPEECH FRAMEWORK FOR U-TSANG, AMDO AND KHAM SPEECH DATASET GENERATION},
  author = {Yutong Liu and Renzeng Duojie and Yuqing Cai and Cheng Huang and Nyima Tashi},
  booktitle = {ICASSP 2026},
  year = {2026}
}