← Search

Yanmin Qian

82 accepted papers

2026

ICASSP 2026 URGENT Speech Enhancement Challenge

ICASSP 2026poster

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, an…

Cited by 0SourcePDFScholar
2026

MEANSE: EFFICIENT GENERATIVE SPEECH ENHANCEMENT WITH MEAN FLOWS

ICASSP 2026poster

Speech enhancement (SE) improves degraded speech's quality, with generative models like flow matching gaining attention for their outstanding perceptual quality. However, the flow-based model requires multiple numbers of function evaluations (NFEs) to achieve stable and satisfactory performance, lea…

Cited by 0SourcePDFScholar
2026

USE: A Unified Model for Universal Sound Separation and Extraction

AAAI 2026technical

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require precisely specified clues to achieve optimal performance. This

Cited by 0SourcePDFScholar
2025

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction

ICASSP 2025accepted

The acoustic background plays a crucial role in natural conversation. It provides context and helps listeners understand the environment, but a strong background makes it difficult for listeners to understand spoken words. The appropriate handling of these backgrounds is situation-dependent: Althoug…

Cited by 0SourceScholar
2025

CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching

NeurIPS 2025poster

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model overlapping speech, and synthesize coherent conversations efficiently…

Cited by 0SourceScholar
2025

Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning

ICASSP 2025accepted

The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability o…

Cited by 0SourceScholar
2025

DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration

NeurIPS 2025poster

Pruning is a widely used technique to compress large language models (LLMs) by removing unimportant weights, but it often suffers from significant performance degradation—especially under semi-structured sparsity constraints. Existing pruning methods primarily focus on estimating the importance of i…

Cited by 0SourcecodeScholar
2025

Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching for Speaker Diarization

ICASSP 2025accepted

Speaker diarization is typically considered as a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore for the first time the use of neural network-based generative methods for speaker diarization. We implement a Flow-Matching (FM) based…

Cited by 0SourceScholar
2025

Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation

ICASSP 2025accepted

Advances in speech synthesis technologies, like text-to-speech (TTS) and voice conversion (VC), have made detecting deepfake speech increasingly challenging. Spoofing countermeasures often struggle to generalize effectively, particularly when faced with unseen attacks. To address this, we propose a…

Cited by 0SourceScholar
2025

SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation

ICASSP 2025accepted

Recently, "textless" speech language models (SLMs) based on speech units have made huge progress in generating naturalistic speech, including non-verbal vocalizations. However, the generated speech samples often lack semantic coherence. In this paper, we propose SLM and LLM Integration for spontaneo…

Cited by 0SourceScholar
2025

SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation

NeurIPS 2025poster

Simultaneous Speech Translation (SimulST) enables real-time cross-lingual communication by jointly optimizing speech recognition and machine translation under strict latency constraints. Existing systems struggle to balance translation quality, latency, and semantic coherence, particularly in multil…

Cited by 0SourceScholar
2025

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

ACL 2025long

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited in scale and diversity, making it challenging to train mode…

2024

CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations

NeurIPS 2024poster

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge. In this paper, we introduce CoVoMix: Conver…

2024

Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection

ICASSP 2024accepted

Machine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages…

Cited by 0SourceScholar
2024

Generation-Based Target Speech Extraction with Speech Discretization and Vocoder

ICASSP 2024accepted

Target speech extraction (TSE) is a task aiming at isolating the speech of a specific target speaker from an audio mixture, with the help of an auxiliary recording of that target speaker. Most existing TSE methods employ discrimination-based models to estimate the target speaker’s proportion in the…

Cited by 0SourceScholar
2024

InstructME: An Instruction Guided Music Edit Framework with Latent Diffusion Models

IJCAI 2024poster

Music editing primarily entails the modification of instrument tracks or remixing in the whole, which offers a novel reinterpretation of the original piece through a series of operations. These music processing methods hold immense potential across various applications but demand substantial experti…

2024

Leveraging in-the-wild Data for Effective Self-supervised Pretraining in Speaker Recognition

ICASSP 2024accepted

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. H…

Cited by 0SourceScholar
2024

Prompt-Driven Target Speech Diarization

ICASSP 2024accepted

We introduce a novel task named ‘target speech diarization’, which seeks to determine ‘when target event occurred’ within an audio signal. We devise a neural architecture called Prompt-driven Target Speech Diarization (PTSD), that works with diverse prompts that specify the target speech events of i…

Cited by 0SourceScholar
2024

Robust Cross-Domain Speaker Verification with Multi-Level Domain Adapters

ICASSP 2024accepted

Speaker verification encounters significant challenges when confronted with diverse domain data, often resulting in performance degradation due to domain mismatch. To enhance performance in cross-domain scenarios, we introduce the Domain Adapter, an adaptable module designed for specific domains. Th…

Cited by 0SourceScholar
2024

TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation

NeurIPS 2024poster

There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models, i.e., a pipeline framework by concatenating speech recognition,…

2023

Code-Switching Text Generation and Injection in Mandarin-English ASR

ICASSP 2023accepted

Code-switching speech refers to a means of expression by mixing two or more languages within a single utterance. Automatic Speech Recognition (ASR) with End-to-End (E2E) modeling for such speech can be a challenging task due to the lack of data. In this study, we investigate text generation and inje…

Cited by 0SourceScholar
2023

ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text Translation

NeurIPS 2023poster

Joint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language. We present ComSL, a speech-language model built atop a composite architecture of public pre-trained speech-only and language-only mode…

2023

Factorized AED: Factorized Attention-Based Encoder-Decoder for Text-Only Domain Adaptive ASR

ICASSP 2023accepted

End-to-end automatic speech recognition (ASR) systems have gained popularity given their simplified architecture and promising results. However, text-only domain adaptation remains a big challenge for E2E systems. Text-to-speech (TTS) based approaches fine-tune ASR models by synthesized speech with…

Cited by 0SourceScholar
2023

HuBERT-AGG: Aggregated Representation Distillation of Hidden-Unit Bert for Robust Speech Recognition

ICASSP 2023accepted

Self-supervised learning (SSL) has attracted widespread research interest since many successful SSL approaches such as wav2vec 2.0 and Hidden-unit BERT (HuBERT) have achieved promising results on speech-related tasks such as automatic speech recognition (ASR). However, few works have been conducted…

Cited by 0SourceScholar
2023

Joint Discriminator and Transfer Based Fast Domain Adaptation For End-To-End Speech Recognition

ICASSP 2023accepted

Adapting End-to-End (E2E) models to unseen domains is still a big challenge since training E2E models requires lots of paired audio and text training data. We propose a novel domain adaptation framework for the E2E model, which only uses the text of the target domain. Moreover, the proposed methods…

Cited by 0SourceScholar
2023

LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer

ICASSP 2023accepted

Traditional automatic speech recognition (ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is more practical in real scenarios. Simply attending longer transcription history for a vanilla neural transducer model shows…

Cited by 0SourceScholar
2023

Multi-Speaker End-to-End Multi-Modal Speaker Diarization System for the MISP 2022 Challenge

ICASSP 2023accepted

This paper presents the design and implementation of our system for Track 1 of the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. We design an end-to-end transformer-based multi-talker system. The transformer backbone is well-suited to capture long-term features, which is cru…

Cited by 0SourceScholar
2023

Predictive Skim: Contrastive Predictive Coding for Low-Latency Online Speech Separation

ICASSP 2023accepted

In online speech separation, there is a trade-off between inherent latency and speech separation performance. When processing the current input audio, looking ahead to more future context usually brings better speech separation performance but increases the algorithm latency, and vice versa. In the…

Cited by 0SourceScholar
2023

Target Sound Extraction with Variable Cross-Modality Clues

ICASSP 2023accepted

Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture of sources. It often uses a model conditioned on a fixed form of target sound clues, such as a sound class label, which l…

Cited by 0SourceScholar
2023

Wespeaker: A Research and Production Oriented Speaker Embedding Learning Toolkit

ICASSP 2023accepted

Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research and production oriented speaker embedding learning toolkit,…

Cited by 0SourceScholar
2022

Exploring Effective Data Utilization for Low-Resource Speech Recognition

ICASSP 2022accepted

Automatic speech recognition (ASR) has suffered great performance degradation when facing low-resource languages with limited training data. In this work, we propose a series of training strategies to exploring more effective data utilization for low-resource speech recognition. In low-resource scen…

Cited by 0SourceScholar
2022

Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification

ICASSP 2022accepted

The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we explore the limits of speech representations learned by different…

Cited by 0SourceScholar
2022

MLP-SVNET: A Multi-Layer Perceptrons Based Network for Speaker Verification

ICASSP 2022accepted

Convolution and self-attention based neural networks have both obtained excellent performance in automatic speaker verification. However, the convolution model often lacks the ability of long-term dependency modeling due to the limitation of receptive field, while the self-attention model is insuffi…

Cited by 0SourceScholar
2022

Optimizing Alignment of Speech and Language Latent Spaces for End-To-End Speech Recognition and Understanding

ICASSP 2022accepted

The advances in attention-based encoder-decoder (AED) networks have brought great progress to end-to-end (E2E) automatic speech recognition (ASR). One way to further improve the performance of AED-based E2E ASR is to introduce an extra text encoder for leveraging extensive text data and thus capture…

Cited by 0SourceScholar
2022

Self-Knowledge Distillation via Feature Enhancement for Speaker Verification

ICASSP 2022accepted

As the most widely used technique, deep speaker embedding learning has become predominant in speaker verification task recently. Very large neural networks such as ECAPA-TDNN and ResNet can achieve the state-of-the-art performance. However, large models are computationally unfriendly in general, whi…

Cited by 0SourceScholar
2022

Skim: Skipping Memory Lstm for Low-Latency Real-Time Continuous Speech Separation

ICASSP 2022accepted

Continuous speech separation for meeting pre-processing has recently become a focused research topic. Compared to the data in utterance-level speech separation, the meeting-style audio stream lasts longer, has an uncertain number of speakers. We adopt the time-domain speech separation method and the…

Cited by 0SourceScholar
2022

Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

ICASSP 2022accepted

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic spe…

Cited by 0SourceScholar
2022

The Sjtu System For Multimodal Information Based Speech Processing Challenge 2021

ICASSP 2022accepted

This paper describes the SJTU system for ICASSP Multi-modal Information based Speech Processing Challenge (MISP) 2021. To solve the speech recognition problem in real complex environments where time-synchronized near- and far-field signals are available for training an enhancement frontend. We build…

Cited by 0SourceScholar
2022

Time-Domain Audio-Visual Speech Separation on Low Quality Videos

ICASSP 2022accepted

Incorporating visual information is a promising approach to improve the performance of speech separation. Many related works have been conducted and provide inspiring results. However, low quality videos appear commonly in real scenarios, which may significantly degrade the performance of normal aud…

Cited by 0SourceScholar
2021

AISpeech-SJTU ASR System for the Accented English Speech Recognition Challenge

ICASSP 2021accepted

This paper describes the AISpeech-SJTU ASR system for the Interspeech-2020 Accented English Speech Recognition Challenge (AESRC). This task is challenging due to the diversity of pronunciation accuracy, intonation speed and pronunciation of some syllables. All participants were restricted to develop…

Cited by 0SourceScholar
2021

AISpeech-SJTU Accent Identification System for the Accented English Speech Recognition Challenge

ICASSP 2021accepted

This paper describes the AISpeech-SJTU system for the accent identification track of the Interspeech-2020 Accented English Speech Recognition Challenge. In this challenge track, only 160-hour accented English data collected from 8 countries and the auxiliary Librispeech dataset are provided for trai…

Cited by 0SourceScholar
2021

Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech Separation

ICASSP 2021accepted

Time-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech enhancement and source separation. Here, we propose to combine ne…

Cited by 0SourceScholar
2021

Dual-Path Modeling for Long Recording Speech Separation in Meetings

ICASSP 2021accepted

The continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straightforward extension of conventional utterance-level speech separation to the CSS task is to segment the long recording wit…

Cited by 0SourceScholar
2021

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

ICASSP 2021accepted

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both singlechannel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performan…

Cited by 0SourceScholar
2021

Self-Supervised Learning Based Domain Adaptation for Robust Speaker Verification

ICASSP 2021accepted

Large performance degradation is often observed for speaker verification systems when applied to a new domain dataset. Given an unlabeled target-domain dataset, unsupervised domain adaptation (UDA) methods, which usually leverage adversarial training strategies, are commonly used to bridge the perfo…

Cited by 0SourceScholar
2021

SynAug: Synthesis-Based Data Augmentation for Text-Dependent Speaker Verification

ICASSP 2021accepted

Text-dependent speaker verification systems trained on large amount of labelled data exhibit remarkable performance. However, collecting the speech from a lot of speakers with target transcript is a lengthy and expensive process. In this work, we propose a synthesis based data augmentation method (S…

Cited by 0SourceScholar
2021

The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and Methods

ICASSP 2021accepted

The variety of accents has posed a big challenge to speech recognition. The Accented English Speech Recognition Challenge (AESRC2020) is designed for providing a common testbed and promoting accent-related research. Two tracks are set in the challenge – English accent recognition (track 1) and accen…

Cited by 0SourceScholar
2021

Towards Data Selection on TTS Data for Children's Speech Recognition

ICASSP 2021accepted

Although great progress has been made on automatic speech recognition (ASR) systems, children’s speech recognition still remains a challenging task. General ASR systems for children’s speech suffer from the lack of corpora and mismatch between children’s and adults’ speech. Efforts have been made to…

Cited by 0SourceScholar
2021

Unit Selection Synthesis Based Data Augmentation for Fixed Phrase Speaker Verification

ICASSP 2021accepted

Data augmentation is commonly used to help build a robust speaker verification system, especially in limited-resource case. However, conventional data augmentation methods usually focus on the diversity of acoustic environment, leaving the lexicon variation neglected. For text dependent speaker veri…

Cited by 0SourceScholar
2020

Channel Invariant Speaker Embedding Learning with Joint Multi-Task and Adversarial Training

ICASSP 2020accepted

Using deep neural network to extract speaker embedding has significantly improved the speaker verification task. However, such embeddings are still vulnerable to channel variability. Previous works have used adversarial training to suppress channel information to extract channel-invariant embedding…

Cited by 0SourceScholar
2020

End-To-End Multi-Speaker Speech Recognition With Transformer

ICASSP 2020accepted

Recently, fully recurrent neural network (RNN) based end-to-end models have been proven to be effective for multi-speaker speech recognition in both the single-channel and multi-channel scenarios. In this work, we explore the use of Transformer models for these tasks by focusing on two aspects. Firs…

Cited by 0SourceScholar
2020

Text Adaptation for Speaker Verification with Speaker-Text Factorized Embeddings

ICASSP 2020accepted

Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system performance. Although this problem can be solved by carefully collecting data with the target speech content, such data c…

Cited by 0SourceScholar
2019

End-to-end Monaural Multi-speaker ASR System without Pretraining

ICASSP 2019accepted

Recently, end-to-end models have become a popular approach as an alternative to traditional hybrid models in automatic speech recognition (ASR). The multi-speaker speech separation and recognition task is a central task in cocktail party problem. In this paper, we present a state-of-the-art monaural…

Cited by 0SourceScholar
2019

Knowledge Distillation for Small Foot-print Deep Speaker Embedding

ICASSP 2019accepted

Deep speaker embedding learning is an effective method for speaker identity modelling. Very deep models such as ResNet can achieve remarkable results but are usually too computationally expensive for real applications with limited resources. On the other hand, simply reducing model size is likely to…

Cited by 0SourceScholar
2018

Adaptive Permutation Invariant Training with Auxiliary Information for Monaural Multi-Talker Speech Recognition

ICASSP 2018accepted

In this paper, we extend our previous work on direct recognition of single-channel multi-talker mixed speech using permutation invariant training (PIT). We propose to adapt the PIT models with auxiliary features such as pitch and i-vector, and to exploit the gender information with multi-task learni…

Cited by 0SourceScholar
2018

Focal Kl-Divergence Based Dilated Convolutional Neural Networks for Co-Channel Speaker Identification

ICASSP 2018accepted

Recognizing the identities of multiple talkers via their overlapped speech is a challenging task, it is also one main difficulty for the “cocktail party problem”. In this paper, a novel dilated convolutional neural network with a focal KL-divergence loss function is proposed to tackle this problem.…

Cited by 0SourceScholar
2018

Generative Adversarial Networks Based Data Augmentation for Noise Robust Speech Recognition

ICASSP 2018accepted

Data augmentation is an effective method to increase the size of training data and reduce the mismatch between training and testing for noise robust speech recognition. Different from the traditional approaches by directly adding noise to the original waveform, in this work we utilize generative adv…

Cited by 0SourceScholar
2018

Joint I-Vector with End-to-End System for Short Duration Text-Independent Speaker Verification

ICASSP 2018accepted

Factor analysis based i-vector has been the state-of-the-art method for speaker verification. Recently, researchers propose to build DNN based end-to-end speaker verification systems and achieve comparable performance with <i xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.…

Cited by 0SourceScholar
2018

Knowledge Transfer in Permutation Invariant Training for Single-Channel Multi-Talker Speech Recognition

ICASSP 2018accepted

This paper proposes a framework that combines teacher-student training and permutation invariant training (PIT) for single-channel multi-talker speech recognition. In contrast to most of conventional teacher-student training methods that aim at compressing the model, the proposed method distills kno…

Cited by 0SourceScholar
2018

Robust Mask Estimation By Integrating Neural Network-Based and Clustering-Based Approaches for Adaptive Acoustic Beamforming

ICASSP 2018accepted

Recently the mask-based beamforming approach received tremendous interest and is widely studied for multi-channel noise robust automatic speech recognition (ASR). Among the known mask estimation models, the neural network based mask estimation approach has received the most attention, resulting in a…

Cited by 0SourceScholar
2016

Improved DNN-based segmentation for multi-genre broadcast audio

ICASSP 2016accepted

Automatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for…

Cited by 0SourceScholar
2016

Integrated adaptation with multi-factor joint-learning for far-field speech recognition

ICASSP 2016accepted

Although great progress has been made in automatic speech recognition (ASR), significant performance degradation still exists in distant talking scenarios due to significantly lower signal power. In this paper, a novel adaptation framework, named integrated adaptation with multi-factor joint-learnin…

Cited by 0SourceScholar
2016

Joint acoustic factor learning for robust deep neural network based automatic speech recognition

ICASSP 2016accepted

Deep neural networks (DNNs) for acoustic modeling have been shown to provide impressive results on many state-of-the-art automatic speech recognition (ASR) applications. However, DNN performance degrades due to mismatches in training and testing conditions and thus adaptation is necessary. In this p…

Cited by 0SourceScholar
2016

Speaker-aware training of LSTM-RNNS for acoustic modelling

ICASSP 2016accepted

Long Short-Term Memory (LSTM) is a particular type of recurrent neural network (RNN) that can model long term temporal dynamics. Recently it has been shown that LSTM-RNNs can achieve higher recognition accuracy than deep feed-forword neural networks (DNNs) in acoustic modelling. However, speaker ada…

Cited by 0SourceScholar
2015

Recurrent neural network language model with structured word embeddings for speech recognition

ICASSP 2015accepted

Due to effective word context encoding and long-term context preserving, recurrent neural network language model (RNNLM) has attracted great interest by showing better performance over back-off n-gram models and feed-forward neural network language models (FNNLM). However, it still has the difficult…

Cited by 0SourceScholar