← Search

Guinan Li

6 accepted papers

2025

Effective and Efficient Mixed Precision Quantization of Speech Foundation Models

ICASSP 2025accepted

This paper presents a novel mixed-precision quantization approach for speech foundation models that tightly integrates mixed-precision learning and quantized model parameter estimation into one single model compression stage. Experiments conducted on LibriSpeech dataset with fine-tuned wav2vec2.0-ba…

Cited by 7SourceScholar
2024

Enhancing Pre-Trained ASR System Fine-Tuning for Dysarthric Speech Recognition Using Adversarial Data Augmentation

ICASSP 2024accepted

Automatic recognition of dysarthric speech remains a highly challenging task to date. Neuro-motor conditions and co-occurring physical disabilities create difficulty in large-scale data collection for ASR system development. Adapting SSL pre-trained ASR models to limited dysarthric speech via data-i…

Cited by 54SourceScholar
2024

Towards Automatic Data Augmentation for Disordered Speech Recognition

ICASSP 2024accepted

Automatic recognition of disordered speech remains a highly challenging task to date due to data scarcity. This paper presents a reinforcement learning (RL) based on-the-fly data augmentation approach for training state-of-the-art PyChain TDNN and end-to-end Conformer ASR systems on such data. The h…

Cited by 0SourceScholar
2024

Towards High-Performance and Low-Latency Feature-Based Speaker Adaptation of Conformer Speech Recognition Systems

ICASSP 2024accepted

Practical application of model-based speaker adaptation techniques to end-to-end ASR systems is hindered by speaker-level data scarcity and latency in speaker-dependent (SD) parameters update. To this end, data-efficient and low-latency rapid feature-based speaker adaptation approaches are proposed…

Cited by 0SourceScholar
2023

Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition

ICASSP 2023accepted

Automatic recognition of disordered speech remains a highly challenging task to date. The underlying neuro-motor conditions, often compounded with co-occurring physical disabilities, lead to the difficulty in collecting large quantities of impaired speech required for ASR system development. This pa…

Cited by 0SourceScholar
2022

Audio-Visual Multi-Channel Speech Separation, Dereverberation and Recognition

ICASSP 2022accepted

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverberation remains a highly challenging task to date. Motivated by the invariance of v…

Cited by 0SourceScholar