ICASSP 2025accepted0 citations

Voice Conversion via Structural Entropy

Linqin Wang, Zhengtao Yu, Shengxiang Gao, Cunli Mao, Ling Dong, Yuxin Huang

Abstract

Voice conversion (VC) aims to transform a person’s voice to resemble that of another person while maintaining the original linguistic content. Existing methods suffer from the blurring of speech representations and the leakage of prosody information. To address this issue, this study introduces SEVC, a novel neural structural entropy-based VC framework. First, we extract self-supervised representations from both the source and reference speech. The representations of the reference speech are structured as a graph. Using two-dimensional (2D) structural entropy (SE), semantically similar representations are clustered together. For the speaker conversion process, each frame of the source speech is treated as a new node, and the most appropriate semantic cluster for each node is identified using SE. Each frame of the source representation is then replaced by the center representation of its corresponding semantic cluster from the reference speech. Finally, a pretrained vocoder synthesizes audio from the transformed representations. Quantitative and qualitative evaluations demonstrate that SEVC improves speaker similarity while maintaining intelligibility scores comparable to existing methods. Additionally, experimental results indicate that SEVC effectively disentangles paralinguistic features from the reference speech, allowing the model to better focus on the semantic content.

BibTeX
@inproceedings{icassp2025_voiceconversionv,
  title = {Voice Conversion via Structural Entropy},
  author = {Linqin Wang and Zhengtao Yu and Shengxiang Gao and Cunli Mao and Ling Dong and Yuxin Huang},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Voice Conversion via Structural Entropy · ICASSP 2025