← Search

Yishuang Li

3 accepted papers

2024

MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis

AAAI 2024technical

The style transfer task in Text-to-Speech (TTS) refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on fixed emotional labels or reference speech clips, whi…

2024

Multivariate Fourier Distribution Perturbation: Domain Shifts with Uncertainty in Frequency Domain

ICASSP 2024accepted

Diversifying training data techniques have achieved tremendous success in Domain Generalization (DG) tasks. The key to diversifying domain data is by increasing the types of domain styles. After investigating this issue from the perspective of the Fourier transform, the domain cue is found to be imp…

Cited by 0SourceScholar
2024

SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification

ICASSP 2024accepted

Recently, pre-trained models (PTMs) have been extensively applied in speaker verification (SV) and greatly boosted system performance. However, mainstream PTMs currently concentrate on using frame-level universal representations. In this paper, we propose a novel pre-training framework that jointly…

Cited by 0SourceScholar