← Search

Danwei Cai

6 accepted papers

2024

Joint Inference of Speaker Diarization and ASR with Multi-Stage Information Sharing

ICASSP 2024accepted

In this paper, we introduce a novel approach that unifies Automatic Speech Recognition (ASR) and speaker diarization in a cohesive framework. Utilizing the synergies between the two tasks, our method effectively extracts speaker-specific information from the lower layers of a pretrained Conformer-ba…

Cited by 0SourceScholar
2023

Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification Systems

ICASSP 2023accepted

An automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system could manipulate a person’s speech signal to make it sound like another speaker’s voice and deceive the speaker verification system. Most countermeasures for voice conv…

Cited by 0SourceScholar
2020

Within-Sample Variability-Invariant Loss for Robust Speaker Recognition Under Noisy Environments

ICASSP 2020accepted

Despite the significant improvements in speaker recognition enabled by deep neural networks, unsatisfactory performance persists under noisy environments. In this paper, we train the speaker embedding network to learn the "clean" embedding of the noisy utterance. Specifically, the network is trained…

Cited by 0SourceScholar
2019

Utterance-level End-to-end Language Identification Using Attention-based CNN-BLSTM

ICASSP 2019accepted

In this paper, we present an end-to-end language identification framework, the attention-based Convolutional Neural Network-Bidirectional Long-short Term Memory (CNN-BLSTM). The model is performed on the utterance level, which means the utterance-level decision can be directly obtained from the outp…

Cited by 0SourceScholar