← Search

Yuzong Liu

7 accepted papers

2024

On-Device Constrained Self-Supervised Learning for Keyword Spotting via Quantization Aware Pre-Training and Fine-Tuning

ICASSP 2024accepted

Large self-supervised models have excelled in various speech processing tasks, but their deployment on resource-limited devices is often impractical due to their substantial memory footprint. Previous studies have demonstrated the effectiveness of self-supervised pre-training for keyword spotting, e…

Cited by 0SourceScholar
2023

Fixed-Point Quantization Aware Training for on-Device Keyword-Spotting

ICASSP 2023accepted

Fixed-point (FXP) inference has proven suitable for embedded devices with limited computational resources, and yet model training is continually performed in floating-point (FLP). FXP training has not been fully explored and the non-trivial conversion from FLP to FXP presents unavoidable performance…

Cited by 0SourceScholar
2023

Self-Supervised Speech Representation Learning for Keyword-Spotting With Light-Weight Transformers

ICASSP 2023accepted

Self-supervised speech representation learning (S3RL) is revolutionizing the way we leverage the ever-growing availability of data. While S3RL related studies typically use large models, we employ light-weight networks to comply with tight memory of compute-constrained devices. We demonstrate the ef…

Cited by 0SourceScholar
2023

Small-Footprint Slimmable Networks for Keyword Spotting

ICASSP 2023accepted

In this work, we present Slimmable Neural Networks applied to the problem of small-footprint keyword spotting. We show that slimmable neural networks allow us to create super-nets from Convolutional Neural Networks and Transformers, from which sub-networks of different sizes can be extracted. We dem…

Cited by 9SourceScholar
2021

Transformer-Transducers for Code-Switched Speech Recognition

ICASSP 2021accepted

We live in a world where 60% of the population can speak two or more languages fluently. Members of these communities constantly switch between languages when having a conversation. As automatic speech recognition (ASR) systems are being deployed to the real-world, there is a need for practical syst…

Cited by 0SourceScholar
2020

Deep Contextualized Acoustic Representations for Semi-Supervised Speech Recognition

ICASSP 2020accepted

We propose a novel approach to semi-supervised automatic speech recognition (ASR). We first exploit a large amount of unlabeled audio data via representation learning, where we reconstruct a temporal slice of filterbank features from past and future context frames. The resulting deep contextualized…

Cited by 0SourceScholar
2019

End-to-end Anchored Speech Recognition

ICASSP 2019accepted

Voice-controlled house-hold devices, like Amazon Echo or Google Home, face the problem of performing speech recognition of device-directed speech in the presence of interfering background speech, i.e., background noise and interfering speech from another person or media device in proximity need to b…

Cited by 20SourceScholar