← Search

Youzhi Tu

8 accepted papers

2025

Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification

ICASSP 2025accepted

In recent years, there has been a surge in the use of a pre-trained speech model as a feature extractor for speaker verification (SV). To reduce model complexity, researchers transfer knowledge from a pre-trained model to a lightweight student model, enabling the latter to reach a performance level…

Cited by 0SourceScholar
2025

Grouped Knowledge Distillation with Adaptive Logit Softening for Speaker Recognition

ICASSP 2025accepted

Recent works suggest that decoupling the information of non-target speakers from that of the target speaker in knowledge distillation (KD) and subsequently emphasizing the former can lead to significant performance improvement. However, a well-trained teacher model typically produces almost zero non…

Cited by 0SourceScholar
2024

Promoting Independence of Depression and Speaker Features for Speaker Disentanglement in Speech-Based Depression Detection

ICASSP 2024accepted

Recent studies have demonstrated the effectiveness of speaker disentanglement in mitigating the interference caused by speaker features in speech-based depression detection. However, the inherent entanglement between depression features and speaker features poses challenges to depression detection.…

Cited by 0SourceScholar
2023

Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations

ICML 2023poster

Self-supervised learning (SSL) speech models such as wav2vec and HuBERT have demonstrated state-of-the-art performance on automatic speech recognition (ASR) and proved to be extremely useful in low label-resource settings. However, the success of SSL models has yet to transfer to utterance-level tas…

Cited by 7SourcePDFScholar
2020

Information Maximized Variational Domain Adversarial Learning for Speaker Verification

ICASSP 2020accepted

Domain mismatch is a common problem in speaker verification. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) to reduce domain mismatch by incorporating an InfoVAE into domain adversarial training (DAT). DAT aims to produce speaker discriminative…

Cited by 0SourceScholar
2019

Semi-supervised Nuisance-attribute Networks for Domain Adaptation

ICASSP 2019accepted

How to overcome the training and test data mismatch in speaker verification systems has been a focus of research recently. In this paper, we propose a semi-supervised nuisance attribute network (SNAN) to reduce the domain mismatch in i-vectors and x-vectors. SNANs are based on the idea of nuisance a…

Cited by 0SourceScholar