← Search

Nguyen Thi Thu Trang

4 accepted papers

2025

O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion

EMNLP 2025

Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these factors remains challenging, often leading to information loss

2025

Voice Conversion for Low-Resource Languages via Knowledge Transfer and Domain-Adversarial Training

ICASSP 2025accepted

Voice conversion (VC) aims to transform speech from a source speaker to a target speaker while preserving the original linguistic content. However, existing VC models typically require large annotated datasets containing transcripts and speaker labels, posing significant challenges for low-resource…

Cited by 0SourceScholar
2025

VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition

ICASSP 2025accepted

Recent research in speaker recognition aims to address vulnerabilities due to variations between enrolment and test utterances, particularly in the multi-genre phenomenon where the utterances are in different speech genres. Previous resources for Vietnamese speaker recognition are either limited in…

Cited by 0SourceScholar
2024

A Robust Pitch-Fusion Model for Speech Emotion Recognition in Tonal Languages

ICASSP 2024accepted

Speech Emotion Recognition (SER) is an essential task in spoken language processing, applicable across various domains. While research on SER systems for English datasets is growing rapidly, the reliability of these models for tonal languages remains a significant concern. Therefore, this paper intr…

Cited by 0SourceScholar