Emotion-Preserving Prosody Anonymization Network for Voice Privacy Protection
Jiabei He, Shiwan Zhao, Jiaming Zhou, Haoqin Sun, Hui Wang, Yong Qin
Abstract
Balancing emotion preservation and privacy protection in voice anonymization presents a significant challenge, particularly due to the difficulty of effectively handling prosody, a key feature in speech. While preserving prosodic features in anonymized speech enhances emotional expression, it also increases the risk of leaking speaker information. To address this conflict, we propose a lightweight Emotion-Preserving Prosody Anonymization (EPPA) network, which extracts speaker-independent prosodic features to preserve speech emotion while converting them into another speaker’s style for anonymization. By combining EPPA with timbre cloning for anonymization while retaining speech content, we achieve a more balanced voice conversion. Evaluated using the Voice Privacy Challenge (VPC) 2024 metrics, our proposed EPPA, utilizing the closest center distance (CCD) anonymization strategy, demonstrates strong performance across emotional expression, content clarity, and privacy protection, achieving the highest ranking in both average and weighted ranks compared to the six baseline solutions.
BibTeX
@inproceedings{icassp2025_emotionpreservin,
title = {Emotion-Preserving Prosody Anonymization Network for Voice Privacy Protection},
author = {Jiabei He and Shiwan Zhao and Jiaming Zhou and Haoqin Sun and Hui Wang and Yong Qin},
booktitle = {ICASSP 2025},
year = {2025}
}