← Search

Huaizhen Tang

4 accepted papers

2024

Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval

ICASSP 2024accepted

Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio has the potential ability to well represent content. Besides,…

Cited by 0SourceScholar
2023

Learning Speech Representations with Flexible Hidden Feature Dimensions

ICASSP 2023accepted

Non-parallel many-to-many voice conversion is a kind of style transfer task in speech. Recently, AutoVC has been applied in this field as a popular solution, as it can achieve distribution-matching style transfer by training only the re- construction loss. However, in order to strike a good balance…

Cited by 0SourceScholar
2023

VQ-CL: Learning Disentangled Speech Representations with Contrastive Learning and Vector Quantization

ICASSP 2023accepted

Voice Conversion(VC) refers to converting the voice characteristics of audio to another one as it is said by other people. Recently, more and more studies have focused on disentangle-based VC, which separates the timbre and linguistic content information from an audio signal to effectively achieve V…

Cited by 0SourceScholar
2022

Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive Learning

ICASSP 2022accepted

Voice Conversion(VC) refers to changing the timbre of a speech while retaining the discourse content. Recently, many works have focused on disentangle-based learning techniques to separate the timbre and the linguistic content information from a speech signal. Once successful, voice conversion will…

Cited by 0SourceScholar