← Search

Feifei Xiong

8 accepted papers

2024

AS-pVAD: A Frame-Wise Personalized Voice Activity Detection Network with Attentive Score Loss

ICASSP 2024accepted

We present a lightweight neural network with attentive score loss for frame-wise personalized voice activity detection (i.e., AS-pVAD). Instead of using an external speaker embedding extractor with a large number of parameters, AS-pVAD employs a lightweight internal model to extract the target speak…

Cited by 0SourceScholar
2023

Deep Subband Network for Joint Suppression of Echo, Noise and Reverberation in Real-Time Fullband Speech Communication

ICASSP 2023accepted

This paper presents a deep and lightweight subband neural network which jointly suppresses the common interference in real-time fullband speech communication: echo, noise and reverberation. Preserving the advantages of spectro-temporal subband network (STSubNet) that requires small amount of resourc…

Cited by 0SourceScholar
2020

Exploring Appropriate Acoustic and Language Modelling Choices for Continuous Dysarthric Speech Recognition

ICASSP 2020accepted

There has been much recent interest in building continuous speech recognition systems for people with severe speech impairments, e.g., dysarthria. However, the datasets that are commonly used are typically designed for tasks other than ASR development, or they contain only isolated words. As such, t…

Cited by 0SourceScholar
2020

Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech Recognition

ICASSP 2020accepted

This paper presents an improved transfer learning framework applied to robust personalised speech recognition models for speakers with dysarthria. As the baseline of transfer learning, a state-of-the-art CNN-TDNN-F ASR acoustic model trained solely on source domain data is adapted onto the target do…

Cited by 0SourceScholar
2019

Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech Recognition

ICASSP 2019accepted

Improving the accuracy of personalised speech recognition for speakers with dysarthria is a challenging research field. In this paper, we explore an approach that non-linearly modifies speech tempo to reduce mismatch between typical and atypical speech. Speech tempo analysis at the phonetic level is…

Cited by 0SourceScholar
2017

Combination strategy based on relative performance monitoring for multi-stream reverberant speech recognition

ICASSP 2017accepted

A multi-stream framework with deep neural network (DNN) classifiers is applied to improve automatic speech recognition (ASR) in environments with different reverberation characteristics. We propose a room parameter estimation model to establish a reliable combination strategy which performs on eithe…

Cited by 0SourceScholar
2017

On DNN posterior probability combination in multi-stream speech recognition for reverberant environments

ICASSP 2017accepted

A multi-stream framework with deep neural network (DNN) classifiers has been applied in this paper to improve automatic speech recognition (ASR) performance in environments with different reverberation characteristics. We propose a room parameter estimation model to determine the stream weights for…

Cited by 0SourceScholar
2015

A study on joint beamforming and spectral enhancement for robust speech recognition in reverberant environments

ICASSP 2015accepted

This work evaluates multi-microphone beamforming and single-microphone spectral enhancement strategies to alleviate the reverberation effect for robust automatic speech recognition (ASR) systems in different reverberant environments characterized by different reverberation times T60 and direct-to-re…

Cited by 3SourceScholar