← Search

Zhengyang Chen

18 accepted papers

2025

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction

ICASSP 2025accepted

The acoustic background plays a crucial role in natural conversation. It provides context and helps listeners understand the environment, but a strong background makes it difficult for listeners to understand spoken words. The appropriate handling of these backgrounds is situation-dependent: Althoug…

Cited by 0SourceScholar
2025

Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning

ICASSP 2025accepted

The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability o…

Cited by 0SourceScholar
2025

Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching for Speaker Diarization

ICASSP 2025accepted

Speaker diarization is typically considered as a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore for the first time the use of neural network-based generative methods for speaker diarization. We implement a Flow-Matching (FM) based…

Cited by 0SourceScholar
2024

Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection

ICASSP 2024accepted

Machine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages…

Cited by 0SourceScholar
2024

Leveraging in-the-wild Data for Effective Self-supervised Pretraining in Speaker Recognition

ICASSP 2024accepted

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. H…

Cited by 0SourceScholar
2024

Prompt-Driven Target Speech Diarization

ICASSP 2024accepted

We introduce a novel task named ‘target speech diarization’, which seeks to determine ‘when target event occurred’ within an audio signal. We devise a neural architecture called Prompt-driven Target Speech Diarization (PTSD), that works with diverse prompts that specify the target speech events of i…

Cited by 0SourceScholar
2024

Robust Cross-Domain Speaker Verification with Multi-Level Domain Adapters

ICASSP 2024accepted

Speaker verification encounters significant challenges when confronted with diverse domain data, often resulting in performance degradation due to domain mismatch. To enhance performance in cross-domain scenarios, we introduce the Domain Adapter, an adaptable module designed for specific domains. Th…

Cited by 0SourceScholar
2023

Multi-Speaker End-to-End Multi-Modal Speaker Diarization System for the MISP 2022 Challenge

ICASSP 2023accepted

This paper presents the design and implementation of our system for Track 1 of the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. We design an end-to-end transformer-based multi-talker system. The transformer backbone is well-suited to capture long-term features, which is cru…

Cited by 0SourceScholar
2023

Wespeaker: A Research and Production Oriented Speaker Embedding Learning Toolkit

ICASSP 2023accepted

Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research and production oriented speaker embedding learning toolkit,…

Cited by 0SourceScholar
2022

Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification

ICASSP 2022accepted

The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we explore the limits of speech representations learned by different…

Cited by 0SourceScholar
2022

MLP-SVNET: A Multi-Layer Perceptrons Based Network for Speaker Verification

ICASSP 2022accepted

Convolution and self-attention based neural networks have both obtained excellent performance in automatic speaker verification. However, the convolution model often lacks the ability of long-term dependency modeling due to the limitation of receptive field, while the self-attention model is insuffi…

Cited by 0SourceScholar
2022

Self-Knowledge Distillation via Feature Enhancement for Speaker Verification

ICASSP 2022accepted

As the most widely used technique, deep speaker embedding learning has become predominant in speaker verification task recently. Very large neural networks such as ECAPA-TDNN and ResNet can achieve the state-of-the-art performance. However, large models are computationally unfriendly in general, whi…

Cited by 0SourceScholar
2022

Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-Training

ICASSP 2022accepted

Self-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years have witnessed great successes in applying self-supervised learning in speech recognition, while limited exploration was attemp…

Cited by 0SourceScholar
2021

Self-Supervised Learning Based Domain Adaptation for Robust Speaker Verification

ICASSP 2021accepted

Large performance degradation is often observed for speaker verification systems when applied to a new domain dataset. Given an unlabeled target-domain dataset, unsupervised domain adaptation (UDA) methods, which usually leverage adversarial training strategies, are commonly used to bridge the perfo…

Cited by 0SourceScholar
2020

Channel Invariant Speaker Embedding Learning with Joint Multi-Task and Adversarial Training

ICASSP 2020accepted

Using deep neural network to extract speaker embedding has significantly improved the speaker verification task. However, such embeddings are still vulnerable to channel variability. Previous works have used adversarial training to suppress channel information to extract channel-invariant embedding…

Cited by 0SourceScholar