ICASSP 2025accepted0 citations

Enhancing Zero-Shot Emotional Voice Conversion via Speaker Adaptation and Duration Prediction

Shiyan Wang, Tianhua Qi, Cheng Lu, Zhaojie Luo, Wenming Zheng

Abstract

Zero-shot Emotional Voice Conversion (EVC) aims to transform a speaker’s emotional state to match a target emotion, even for speakers and emotion categories that were not encountered during training, thereby enhancing the generalization ability of traditional EVC systems. Despite advancements in the field, existing methods often face challenges in preserving speaker identity and ensuring the naturalness of emotional expression, particularly in the context of rhythm modeling. To this end, we propose the Zero-Shot Emotion Voice Conversion (ZSEVC) model, which leverages self-supervised learning for speaker adaptation and duration prediction. To adjust speech rhythm in alignment with the target emotional state, we introduce a rhythm-aware content encoder that captures and refines discrete speech units at a finer granularity. Additionally, a hierarchical emotion fusion scheme is employed to integrate emotional features with content features, enhancing both pronunciation accuracy and emotional expressiveness. Moreover, a residual speaker-emotion fusion module is incorporated to better adapt speaker characteristics to emotional prosodic variation. Experimental results show ZSEVC’s superior performance in terms of naturalness and speaker similarity in zero-shot scenario, successfully generating emotional speeches for unseen emotions and speakers. Speech samples are available at https://wosyoo.github.io/ZSEVC.

BibTeX
@inproceedings{icassp2025_enhancingzerosho,
  title = {Enhancing Zero-Shot Emotional Voice Conversion via Speaker Adaptation and Duration Prediction},
  author = {Shiyan Wang and Tianhua Qi and Cheng Lu and Zhaojie Luo and Wenming Zheng},
  booktitle = {ICASSP 2025},
  year = {2025}
}