← Search

Xiaofei Li

24 accepted papers

2026

UMA-SPLIT: UNIMODAL AGGREGATION FOR BOTH ENGLISH AND MANDARIN NON-AUTOREGRESSIVE SPEECH RECOGNITION

ICASSP 2026poster

This paper proposes a unimodal aggregation (UMA) based nonautoregressive model for both English and Mandarin speech recognition. The original UMA explicitly segments and aggregates acoustic frames (with unimodal weights that first monotonically increase and then decrease) of the same text token to l…

Cited by 0SourcePDFScholar
2024

Frame-Wise Streaming end-to-end Speaker Diarization with Non-Autoregressive Self-Attention-Based Attractors

ICASSP 2024accepted

This work proposes a frame-wise online/streaming end-to-end neural diarization (FS-EEND) method in a frame-in-frame-out fashion. To frame-wisely detect a flexible number of speakers and extract/update their corresponding attractors, we propose to leverage a causal speaker embedding encoder and an on…

Cited by 0SourceScholar
2024

RVAE-EM: Generative Speech Dereverberation Based On Recurrent Variational Auto-Encoder And Convolutive Transfer Function

ICASSP 2024accepted

In indoor scenes, reverberation is a crucial factor in degrading the perceived quality and intelligibility of speech. In this work, we propose a generative dereverberation method. Our approach is based on a probabilistic model utilizing a recurrent variational auto-encoder (RVAE) network and the con…

Cited by 0SourceScholar
2024

RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization

NeurIPS 2024poster

The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded datasets. However, the acoustic mismatch between simulated and re…

2023

Locate, Refine and Restore: A Progressive Enhancement Network for Camouflaged Object Detection

IJCAI 2023poster

Camouflaged Object Detection (COD) aims to segment objects that blend in with their surroundings. Most existing methods mainly tackle this issue by a single-stage framework, which tends to degrade performance in the face of small objects, low-contrast objects and objects with diverse appearances. In…

Cited by 31SourcePDFScholar
2023

Natural Language Instruction Understanding for Robotic Manipulation: a Multisensory Perception Approach

ICRA 2023poster

It has always been expected that the robot can understand the natural language instruction and thus a more natural human-robot interaction is achieved. Currently, the robot usually interprets the instruction by visually grounding the textual information to its surroundings, while it may be not enoug…

Cited by 8SourceScholar
2023

What to Fuse and How to Fuse: Exploring Emotion and Personality Fusion Strategies for Explainable Mental Disorder Detection

ACL 2023findings

Mental health disorders (MHD) are increasingly prevalent worldwide and constitute one of the greatest challenges facing our healthcare systems and modern societies in general. In response to this societal challenge, there has been a surge in digital mental health research geared towards the developm…

Cited by 14SourcePDFScholar
2022

Multi-Channel Narrow-Band Deep Speech Separation with Full-Band Permutation Invariant Training

ICASSP 2022accepted

This paper addresses the problem of multi-channel multi-speech separation based on deep learning techniques. In the short time Fourier transform domain, we propose an end-to-end narrow-band network that directly takes as input the multi-channel mixture signals of one frequency, and outputs the separ…

Cited by 0SourceScholar
2022

SRP-DNN: Learning Direct-Path Phase Difference for Multiple Moving Sound Source Localization

ICASSP 2022accepted

Multiple moving sound source localization in real-world scenarios remains a challenging issue due to interaction between sources, time-varying trajectories, distorted spatial cues, etc. In this work, we propose to use deep learning techniques to learn competing and time-varying direct-path phase dif…

Cited by 0SourceScholar
2021

AcousticFusion: Fusing Sound Source Localization to Visual SLAM in Dynamic Environments

IROS 2021poster

Dynamic objects in the environment, such as people and other agents, lead to challenges for existing simultaneous localization and mapping (SLAM) approaches. To deal with dynamic environments, computer vision researchers usually apply some learning-based object detectors to remove these dynamic obje…

Cited by 22SourceScholar
2021

Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement

ICASSP 2021accepted

This paper proposes a full-band and sub-band fusion model, named as FullSubNet, for single-channel real-time speech enhancement. Full-band and sub-band refer to the models that input full-band and sub-band noisy spectral feature, output full-band and sub-band speech target, respectively. The sub-ban…

Cited by 0SourceScholar
2021

Supervised Direct-Path Relative Transfer Function Learning for Binaural Sound Source Localization

ICASSP 2021accepted

Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two channels. Though DP-RTF fully encodes the sound directional cues and serves as a reliable localization feature, it is often erroneously estimated in the presence of noise an…

Cited by 0SourceScholar
2018

Accounting for Room Acoustics in Audio-Visual Multi-Speaker Tracking

ICASSP 2018accepted

Multiple-speaker tracking is a crucial task for many applications. In real-world scenarios, exploiting the complementarity between auditory and visual data enables to track people outside the visual field of view. However, practical methods must be robust to changes in acoustic conditions, e.g. reve…

Cited by 0SourceScholar
2017

Audio source separation based on convolutive transfer function and frequency-domain lasso optimization

ICASSP 2017accepted

This paper addresses the problem of under-determined convolutive audio source separation in a semi-oracle configuration where the mixing filters are assumed to be known. We propose a separation procedure based on the convolutive transfer function (CTF), which is a more appropriate model for strongly…

Cited by 0SourceScholar
2016

Non-stationary noise power spectral density estimation based on regional statistics

ICASSP 2016accepted

Estimating the noise power spectral density (PSD) is essential for single channel speech enhancement algorithms. In this paper, we propose a noise PSD estimation approach based on regional statistics. The proposed regional statistics consist of four features representing the statistics of the past a…

Cited by 0SourceScholar
2016

Reverberant sound localization with a robot head based on direct-path relative transfer function

IROS 2016poster

This paper addresses the problem of sound-source localization (SSL) with a robot head, which remains a challenge in real-world environments. In particular we are interested in locating speech sources, as they are of high interest for human-robot interaction. The microphone-pair response correspondin…

Cited by 44SourceScholar
2015

Estimation of relative transfer function in the presence of stationary noise based on segmental power spectral density matrix subtraction

ICASSP 2015accepted

This paper addresses the problem of relative transfer function (RTF) estimation in the presence of stationary noise. We propose an RTF identification method based on segmental power spectral density (PSD) matrix subtraction. First multiple channel microphone signals are divided into segments corresp…

Cited by 0SourceScholar