← Search

Duc Le

16 accepted papers

2025

Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens

ICASSP 2025accepted

Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transducer-based speech translation model that outputs discrete speech tokens in a low…

Cited by 0SourceScholar
2024

PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding

EMNLP 2024finding

Spoken Language Understanding (SLU) is a critical component of voice assistants; it consists of converting speech to semantic parses for task execution. Previous works have explored end-to-end models to improve the quality and robustness of SLU models with Deliberation, however these models have rem…

Cited by 0SourcePDFScholar
2024

STEMGEN: A Music Generation Model That Listens

ICASSP 2024accepted

End-to-end generation of musical audio using deep learning techniques has seen an explosion of activity recently. However, most models concentrate on generating fully mixed music in response to abstract conditioning information. In this work, we present an alternative paradigm for producing music ge…

Cited by 0SourceScholar
2023

Factorized Blank Thresholding for Improved Runtime Efficiency of Neural Transducers

ICASSP 2023accepted

We show how factoring the RNN-T’s output distribution can significantly reduce the computation cost and power consumption for on-device ASR inference with no loss in accuracy. With the rise in popularity of neural-transducer type models like the RNN-T for on-device ASR, optimizing RNN-T’s runtime ef…

Cited by 0SourceScholar
2023

ICASSP 2023 Spoken Language Understanding Grand Challenge

ICASSP 2023accepted

Spoken language understanding (SLU) is a important field between the Speech and NLP community focused on converting a users’ speech utterance into an executable semantic parse. In order to facilitate open research in this space, we introduce the 1st Spoken Language Understanding challenge hosted at…

Cited by 0SourceScholar
2023

Improving fast-slow Encoder based Transducer with Streaming Deliberation

ICASSP 2023accepted

This paper introduces a fast-slow encoder based transducer with streaming deliberation for end-to-end automatic speech recognition. We aim to improve the recognition accuracy of the fast-slow encoder based transducer while keeping its latency low by integrating a streaming deliberation model. Specif…

Cited by 0SourceScholar
2023

Learning ASR Pathways: A Sparse Multilingual ASR Model

ICASSP 2023accepted

Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in multilingual ASR, language-agnostic pruning may lead to severe performance drops on some languages because language-agnostic pruning masks may not fit all languages and discard important language-spe…

Cited by 0SourceScholar
2023

Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities

ICASSP 2023accepted

End-to-end multilingual ASR has become more appealing because of several reasons such as simplifying the training and deployment process and positive performance transfer from high-resource to low-resource languages. However, scaling up the number of languages, total hours, and number of unique toke…

Cited by 0SourceScholar
2022

Joint Audio/Text Training for Transformer Rescorer of Streaming Speech Recognition

EMNLP 2022finding

Recently, there has been an increasing interest in two-pass streaming end-to-end speech recognition (ASR) that incorporates a 2nd-pass rescoring model on top of the conventional 1st-pass streaming ASR model to improve recognition accuracy while keeping latency low. One of the latest 2nd-pass rescori…

Cited by 8SourcePDFScholar
2022

Neural-FST Class Language Model for End-to-End Speech Recognition

ICASSP 2022accepted

We propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transducers (FSTs) in a mathematically consistent framework. Our method utilizes a background NNLM which models generic backgroun…

Cited by 0SourceScholar
2021

Emformer: Efficient Memory Transformer Based Acoustic Model for Low Latency Streaming Speech Recognition

ICASSP 2021accepted

This paper proposes an efficient memory transformer Emformer for low latency streaming speech recognition. In Emformer, the long-range history context is distilled into an augmented memory bank to reduce self-attention’s computation complexity. A cache mechanism saves the computation for the key and…

Cited by 0SourceScholar
2021

Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer

ICASSP 2021accepted

Recurrent Neural Network Transducer (RNN-T), like most end-to-end speech recognition model architectures, has an implicit neural network language model (NNLM) and cannot easily leverage unpaired text data during training. Previous work has proposed various fusion methods to incorporate external NNLM…

Cited by 0SourceScholar
2020

G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR

ICASSP 2020accepted

Grapheme-based acoustic modeling has recently been shown to outperform phoneme-based approaches in both hybrid and end-to-end automatic speech recognition (ASR), even on non-phonemic languages like English. However, graphemic ASR still has problems with low-frequency words that do not follow the sta…

Cited by 19SourceScholar
2020

Transformer-Based Acoustic Modeling for Hybrid Speech Recognition

ICASSP 2020accepted

We propose and evaluate transformer-based acoustic models (AMs) for hybrid speech recognition. Several modeling choices are discussed in this work, including various positional embedding methods and an iterated loss to enable training deep transformers. We also present a preliminary study of using l…

Cited by 0SourceScholar