← Search

Xiaodan Zhuang

4 accepted papers

2025

Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition

ICASSP 2025accepted

This paper presents an efficient decoding approach for end-to-end automatic speech recognition (E2E-ASR) with large language models (LLMs). Although shallow fusion is the most common approach to incorporate language models into E2E-ASR decoding, we face two practical problems with LLMs. (1) LLM infe…

Cited by 0SourceScholar
2023

Variable Attention Masking for Configurable Transformer Transducer Speech Recognition

ICASSP 2023accepted

This work studies the use of attention masking in transformer transducer based speech recognition for building a single configurable model for different deployment scenarios. We present a comprehensive set of experiments comparing fixed masking, where the same attention mask is applied at every fram…

Cited by 0SourceScholar
2020

SNDCNN: Self-Normalizing Deep CNNs with Scaled Exponential Linear Units for Speech Recognition

ICASSP 2020accepted

Very deep CNNs achieve state-of-the-art results in both computer vision and speech recognition, but are difficult to train. The most popular way to train very deep CNNs is to use shortcut connections (SC) together with batch normalization (BN). Inspired by Self-Normalizing Neural Networks, we propos…

Cited by 41SourceScholar
2019

Exploring Retraining-free Speech Recognition for Intra-sentential Code-switching

ICASSP 2019accepted

Code Switching refers to the phenomenon of changing languages within a sentence or discourse, and it represents a challenge for conventional automatic speech recognition systems deployed to tackle a single target language. The code switching problem is complicated by the lack of multi-lingual traini…

Cited by 0SourceScholar