← Search

Zexin Cai

8 accepted papers

2025

HLTCOE Submission to the VoicePrivacy Attacker Challenge

ICASSP 2025accepted

We describe our submission to the 2024 VoicePrivacy Attacker Challenge. We propose three main categories of methods to improve ASV performance against anonymized speech: improvements to the underlying classifier, alternative distance metrics when computing ASV scores, and kNN-VC normalization. By si…

Cited by 0SourceScholar
2023

Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification Systems

ICASSP 2023accepted

An automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system could manipulate a person’s speech signal to make it sound like another speaker’s voice and deceive the speaker verification system. Most countermeasures for voice conv…

Cited by 0SourceScholar
2022

SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and Machines

ICASSP 2022accepted

Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people’s attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for…

Cited by 0SourceScholar
2018

A Novel Learnable Dictionary Encoding Layer for End-to-End Language Identification

ICASSP 2018accepted

A novel learnable dictionary encoding layer is proposed in this paper for end-to-end language identification. It is inline with the conventional GMM i-vector approach both theoretically and practically. We imitate the mechanism of traditional GMM training and Supervector encoding procedure on the to…

Cited by 79SourceScholar
2018

Insights in-to-End Learning Scheme for Language Identification

ICASSP 2018accepted

A novel interpretable end-to-end learning scheme for language identification is proposed. It is in line with the classical GMM i-vector methods both theoretically and practically. In the end-to-end pipeline, a general encoding layer is employed on top of the frontend CNN, so that it can encode the v…

Cited by 20SourceScholar