← Search

Xiong Xiao

26 accepted papers

2026

FedAlign: Differentially Private Distribution Alignment for Non-IID Federated Learning

CVPR 2026

Federated Learning (FL) enables collaborative model training without sharing raw data, but client data are often Non-Independent and Identically Distributed (Non-IID), which often slow convergence and degrade global performance. Meanwhile, privacy preservation is also a critical concern in FL. To ad

Cited by 0SourceScholar
2026

Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio

ICASSP 2026poster

Joint automatic speech recognition (ASR) and speaker diarization aim to answer the question "who spoke what" in multi-speaker scenarios. In this paper, we present an end-to-end speech large language model (Speech-LLM) for Joint strEamable DIarization and aSr (JEDIS-LLM). The model is trained only on…

Cited by 0SourcePDFScholar
2024

Profile-Error-Tolerant Target-Speaker Voice Activity Detection

ICASSP 2024accepted

Target-Speaker Voice Activity Detection (TS-VAD) utilizes a set of speaker profiles alongside an input audio signal to perform speaker diarization. While its superiority over conventional methods has been demonstrated, the method can suffer from errors in speaker profiles, as those profiles are typi…

Cited by 0SourceScholar
2023

Target Speaker Voice Activity Detection with Transformers and Its Integration with End-To-End Neural Diarization

ICASSP 2023accepted

This paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model’s drawback of being unable to handle an arbitrary number of speakers, we investigate model architectures that use input tensors with va…

Cited by 0SourceScholar
2022

Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers Using End-to-End Speaker-Attributed ASR

ICASSP 2022accepted

This paper presents Transcribe-to-Diarize, a new approach for neural speaker diarization that uses an end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR). The E2E SA-ASR is a joint model that was recently proposed for speaker counting, multi-talker speech recognition, and speak…

Cited by 0SourceScholar
2021

Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020

ICASSP 2021accepted

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2020. We will first explain our system design to address issues in handling real multi-talker reco…

Cited by 0SourceScholar
2020

Continuous Speech Separation: Dataset and Analysis

ICASSP 2020accepted

This paper describes a dataset and protocols for evaluating continuous speech separation algorithms. Most prior speech separation studies use pre-segmented audio signals, which are typically generated by mixing speech utterances on computers so that they fully overlap. Also, the separation algorithm…

Cited by 0SourceScholar
2020

Speaker Diarization with Session-Level Speaker Embedding Refinement Using Graph Neural Networks

ICASSP 2020accepted

Deep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a global loss defined on the training data, which could be suboptimal for distinguishing speakers locally in a specific meet…

Cited by 0SourceScholar
2019

Low-latency Speaker-independent Continuous Speech Separation

ICASSP 2019accepted

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of which contains no overlapping speech segment. A separated, or cleaned, version of e…

Cited by 0SourceScholar
2019

Single-channel Speech Extraction Using Speaker Inventory and Attention Network

ICASSP 2019accepted

Neural network-based speech separation has received a surge of interest in recent years. Previously proposed methods either are speaker independent or extract a target speaker's voice by using his or her voice snippet. In applications such as home devices or office meeting transcriptions, a possible…

Cited by 76SourceScholar
2018

Developing Far-Field Speaker System Via Teacher-Student Learning

ICASSP 2018accepted

In this study, we develop the keyword spotting (KWS) and acoustic model (AM) components in a far-field speaker system. Specifically, we use teacher-student (T/S) learning to adapt a close-talk well-trained production AM to far-field by using parallel close-talk and simulated far-field data. We also…

Cited by 0SourceScholar
2018

Efficient Integration of Fixed Beamformers and Speech Separation Networks for Multi-Channel Far-Field Speech Separation

ICASSP 2018accepted

Speech separation research has significantly progressed in recent years thanks to the rapid advances in deep learning technology. However the performance of recently proposed single-channel neural network-based speech separation methods is still limited especially in reverberant environments. To pus…

Cited by 0SourceScholar
2018

Single Channel Speech Separation with Constrained Utterance Level Permutation Invariant Training Using Grid LSTM

ICASSP 2018accepted

Utterance level permutation invariant training (uPIT) technique is a state-of-the-art deep learning architecture for speaker independent multi-talker separation. uPIT solves the label ambiguity problem by minimizing the mean square error (MSE) over all permutations between outputs and targets. Howev…

Cited by 0SourceScholar
2017

On time-frequency mask estimation for MVDR beamforming with application in robust speech recognition

ICASSP 2017accepted

Acoustic beamforming has played a key role in the robust automatic speech recognition (ASR) applications. Accurate estimates of the speech and noise spatial covariance matrices (SCM) are crucial for successfully applying the minimum variance distortionless response (MVDR) beamforming. Reliable estim…

Cited by 0SourceScholar
2016

An expectation-maximization eigenvector clustering approach to direction of arrival estimation of multiple speech sources

ICASSP 2016accepted

This paper presents an eigenvector clustering approach for estimating the direction of arrival (DOA) of multiple speech signals using a microphone array. Existing clustering approaches usually only use low frequencies to avoid spatial aliasing. In this study, we propose a probabilistic eigenvector c…

Cited by 0SourceScholar
2016

Approximate search of audio queries by using DTW with phone time boundary and data augmentation

ICASSP 2016accepted

Dynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DT…

Cited by 0SourceScholar
2016

Deep beamforming networks for multi-channel speech recognition

ICASSP 2016accepted

Despite the significant progress in speech recognition enabled by deep neural networks, poor performance persists in some scenarios. In this work, we focus on far-field speech recognition which remains challenging due to high levels of noise and reverberation in the captured speech signals. We propo…

Cited by 0SourceScholar
2016

Exemplar-inspired strategies for low-resource spoken keyword search in Swahili

ICASSP 2016accepted

We present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational…

Cited by 0SourceScholar
2016

Keyword search using query expansion for graph-based rescoring of hypothesized detections

ICASSP 2016accepted

In this work, we propose a novel framework for rescoring keyword search (KWS) detections using acoustic samples extracted from the training data. We view the keyword rescoring task as an information retrieval task and adopt the idea of query expansion. We expand a textual keyword with multiple speec…

Cited by 0SourceScholar
2016

Speaker-aware training of LSTM-RNNS for acoustic modelling

ICASSP 2016accepted

Long Short-Term Memory (LSTM) is a particular type of recurrent neural network (RNN) that can model long term temporal dynamics. Recently it has been shown that LSTM-RNNs can achieve higher recognition accuracy than deep feed-forword neural networks (DNNs) in acoustic modelling. However, speaker ada…

Cited by 0SourceScholar
2016

Spoofing detection from a feature representation perspective

ICASSP 2016accepted

Spoofing detection, which discriminates the spoofed speech from the natural speech, has gained much attention recently. Low-dimensional features that are used in speaker recognition/verification are also used in spoofing detection. Unfortunately, they don't capture sufficient information required fo…

Cited by 0SourceScholar
2015

A learning-based approach to direction of arrival estimation in noisy and reverberant environments

ICASSP 2015accepted

This paper presents a learning-based approach to the task of direction of arrival estimation (DOA) from microphone array input. Traditional signal processing methods such as the classic least square (LS) method rely on strong assumptions on signal models and accurate estimations of time delay of arr…

Cited by 0SourceScholar
2015

Language independent query-by-example spoken term detection using N-best phone sequences and partial matching

ICASSP 2015accepted

In this paper, we propose a partial sequence matching based symbolic search (SS) method for the task of language independent query-by-example spoken term detection. One main drawback of conventional SS approach is the high miss rate for long queries. This is due to high variations in symbol represen…

Cited by 0SourceScholar
2015

Low-resource keyword search strategies for tamil

ICASSP 2015accepted

We propose strategies for a state-of-the-art keyword search (KWS) system developed by the SINGA team in the context of the 2014 NIST Open Keyword Search Evaluation (OpenKWS14) using conversational Tamil provided by the IARPA Babel program. To tackle low-resource challenges and the rich morphological…

Cited by 0SourceScholar