← Search

Xiaoyi Qin

8 accepted papers

2024

Multi-Objective Progressive Clustering for Semi-Supervised Domain Adaptation in Speaker Verification

ICASSP 2024accepted

Utilizing the pseudo-labeling algorithm with large-scale unlabeled data becomes crucial for semi-supervised domain adaptation in speaker verification tasks. In this paper, we propose a novel pseudo-labeling method named Multi-objective Progressive Clustering (MoPC), specifically designed for semi-su…

Cited by 0SourceScholar
2024

Voxblink: A Large Scale Speaker Verification Dataset on Camera

ICASSP 2024accepted

In this paper, we introduce a large-scale and high-quality audiovisual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature o…

Cited by 0SourceScholar
2023

Target-Speaker Voice Activity Detection Via Sequence-to-Sequence Prediction

ICASSP 2023accepted

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD) method that can efficiently address the joint modeling of large…

Cited by 0SourceScholar
2022

Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for the M2met Challenge

ICASSP 2022accepted

DukeECE. As the highly overlapped speech exists in the dataset, we employ an x-vector-based target-speaker voice activity detection (TS-VAD) to find the overlap between speakers. Firstly, we separately train a single-channel model for each of the 8 channels and fuse the results. In addition, we also…

Cited by 32SourceScholar
2022

SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and Machines

ICASSP 2022accepted

Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people’s attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for…

Cited by 0SourceScholar
2022

Simple Attention Module Based Speaker Verification with Iterative Noisy Label Detection

ICASSP 2022accepted

Recently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speaker verification system. This paper introduces an alternative effective yet simple one, i.e., simple attention module (Si…

Cited by 0SourceScholar
2022

Towards Lightweight Applications: Asymmetric Enroll-Verify Structure for Speaker Verification

ICASSP 2022accepted

With the development of deep learning, automatic speaker verification has made considerable progress over the past few years. However, to design a lightweight and robust system with limited computational resources is still a challenging problem. Traditionally, a speaker verification system is symmet…

Cited by 0SourceScholar