← Search

Xiao-Lei Zhang

14 accepted papers

2026

HAMLET: Hyperadaptive Agent-based Modeling for Live Embodied Theatrics

ICLR 2026poster

Creating an immersive and interactive theatrical experience is a long-term goal in the field of interactive narrative. The emergence of large language model (LLM) is providing a new path to achieve this goal. However, existing LLM-based drama generation methods often result in agents that lack initi…

Cited by 0SourcecodeScholar
2025

Co-Attention Based Multi-Channel TF-GridNet for Speech Separation with Ad-Hoc Microphone Arrays

ICASSP 2025accepted

Speech separation using ad-hoc microphone arrays has been explored, but there is still significant room for improvement, especially in complex scenarios with varying channel conditions. Co-attention, a feature fusion mechanism, is widely used in multimodal fusion to capture the cooperation between m…

Cited by 0SourceScholar
2024

Exploiting A Quantum Multiple Kernel Learning Approach For Low-Resource Spoken Command Recognition

ICASSP 2024accepted

We propose a theoretical analysis of quantum projection learning (QPL) that employs multiple kernels, highlighting its advantages through representation error analysis. Building upon previous studies that utilized a single quantum kernel-based method, we further investigate a quantum projection fram…

Cited by 0SourceScholar
2023

Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames

ICASSP 2023accepted

Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The cor…

Cited by 0SourceScholar
2023

Optimizing Quantum Federated Learning Based on Federated Quantum Natural Gradient Descent

ICASSP 2023accepted

Quantum federated learning (QFL) is a quantum extension of the classical federated learning model across multiple local quantum devices. An efficient optimization algorithm is always expected to minimize the communication overhead among different quantum participants. In this work, we propose an eff…

Cited by 0SourceScholar
2023

Soft Label Coding for end-to-end Sound Source Localization with ad-hoc Microphone Arrays

ICASSP 2023accepted

Recently, an end-to-end two-dimensional sound source localization algorithm with ad-hoc microphone arrays formulates the sound source localization problem as a classification problem. The algorithm divides the target indoor space into a set of local areas, and predicts the local area where the speak…

Cited by 0SourceScholar
2023

Wekws: A Production First Small-Footprint End-to-End Keyword Spotting Toolkit

ICASSP 2023accepted

Keyword spotting (KWS) enables speech-based user interaction and gradually becomes an indispensable component of smart devices. Recently, end-to-end (E2E) methods have be-come the most popular approach for on-device KWS tasks. However, there is still a gap between the research and deployment of E2E…

Cited by 0SourceScholar
2022

End-To-End Multi-Modal Speech Recognition with Air and Bone Conducted Speech

ICASSP 2022accepted

Improving the performance of automatic speech recognition (ASR) in adverse acoustic environments is a long-term tough task. Although many robust ASR systems based on conventional microphones have been developed, their performance with air-conducted (AC) speech is still far from satisfactory in low s…

Cited by 0SourceScholar
2021

Transformer-Based End-to-End Speech Recognition with Local Dense Synthesizer Attention

ICASSP 2021accepted

Recently, several studies reported that dot-product self-attention (SA) may not be indispensable to the state-of-the-art Transformer models. Motivated by the fact that dense synthesizer attention (DSA), which dispenses with dot products and pairwise interactions, achieved competitive results in many…

Cited by 0SourceScholar
2020

Partial AUC Optimization Based Deep Speaker Embeddings with Class-Center Learning for Text-Independent Speaker Verification

ICASSP 2020accepted

Deep embedding based text-independent speaker verification has demonstrated superior performance to traditional methods in many challenging scenarios. Its loss functions can be generally categorized into two classes, i.e., verification and identification. The verification loss functions match the pi…

Cited by 0SourceScholar
2019

AUC Optimization for Deep Learning Based Voice Activity Detection

ICASSP 2019accepted

Voice activity detection (VAD) based on deep neural networks (DNN) has demonstrated good performance in adverse acoustic environments. Current DNN based VAD optimizes a surrogate function, e.g. minimum cross-entropy or minimum squared error, at a given decision threshold. However, VAD usually works…

Cited by 0SourceScholar
2019

Robust Sparse Multichannel Active Noise Control

ICASSP 2019accepted

Multichannel active noise control (MC-ANC) aims to cancel low-frequency noise in an enclosure. If noise sources are distributed sparsely in space, adding an ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf> -norm constraint to the standard MC-AN…

Cited by 0SourceScholar