← Search

Weizhong Zhu

8 accepted papers

2022

Speaker Normalization for Self-Supervised Speech Emotion Recognition

ICASSP 2022accepted

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts usually harm a model’s ability to generalize. To address thi…

Cited by 0SourceScholar
2022

Speech Emotion Recognition Using Self-Supervised Features

ICASSP 2022accepted

Self-supervised pre-trained features have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of speech emotion recognition (SER) still need further investigation. In this paper we introduce a modular End-to-End (E2E) SER…

Cited by 0SourceScholar
2022

Towards A Common Speech Analysis Engine

ICASSP 2022accepted

Recent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based systems are not yet considered state-of-the-art.We propose leveraging recent advance…

Cited by 0SourceScholar
2017

Speaker diarization: A perspective on challenges and opportunities from theory to practice

ICASSP 2017accepted

This paper discusses some challenges and opportunities in developing a speaker diarization system for operation on real world call center telephony data. We contrast some of the differences between a standard data set akin to NIST evaluations and those found in call centers. In exploring these diffe…

Cited by 0SourceScholar
2015

Nearest neighbor based i-vector normalization for robust speaker recognition under unseen channel conditions

ICASSP 2015accepted

Many state-of-the-art speaker recognition engines use i-vectors to represent variable-length acoustic signals in a fixed low-dimensional total variability subspace. While such systems perform well under seen channel conditions, their performance greatly degrades under unseen channel scenarios. Accor…

Cited by 0SourceScholar