← Search

Cong-Thanh Do

6 accepted papers

2026

DUAL-SPACE KNOWLEDGE DISTILLATION WITH KEY-QUERY MATCHING FOR LARGE LANGUAGE MODELS WITH VOCABULARY MISMATCH

ICASSP 2026poster

Large language models (LLMs) achieve state-of-the-art (SOTA) performance across language tasks, but are costly to deploy due to their size and resource demands. Knowledge Distillation (KD) addresses this by training smaller Student models to mimic larger Teacher models, improving efficiency without…

Cited by 0SourcePDFScholar
2023

Cumulative Attention Based Streaming Transformer ASR with Internal Language Model Joint Training and Rescoring

ICASSP 2023accepted

This paper presents an approach to improve the performance of streaming Transformer ASR by introducing an internal language model (ILM) as a part of the decoder layers. In the recently pro- posed cumulative attention (CA) based streaming ASR system, only the last or top few decoder layers are equipp…

Cited by 0SourceScholar
2021

Multiple-Hypothesis CTC-Based Semi-Supervised Adaptation of End-to-End Speech Recognition

ICASSP 2021accepted

This paper proposes an adaptation method for end-to-end speech recognition. In this method, multiple automatic speech recognition (ASR) 1-best hypotheses are integrated in the computation of the connectionist temporal classification (CTC) loss function. The integration of multiple ASR hypotheses hel…

Cited by 0SourceScholar
2021

Train Your Classifier First: Cascade Neural Networks Training from Upper Layers to Lower Layers

ICASSP 2021accepted

Although the lower layers of a deep neural network learn features which are transferable across datasets, these layers are not transferable within the same dataset. That is, in general, freezing the trained feature extractor (the lower layers) and retraining the classifier (the upper layers) on the…

Cited by 0SourceScholar
2020

Learning Noise Invariant Features Through Transfer Learning For Robust End-to-End Speech Recognition

ICASSP 2020accepted

End-to-end models yield impressive speech recognition results on clean datasets while having inferior performance on noisy datasets. To address this, we propose transfer learning from a clean dataset (WSJ) to a noisy dataset (CHiME4) for connectionist temporal classification models. We argue that th…

Cited by 19SourceScholar
2019

Subband Temporal Envelope Features and Data Augmentation for End-to-end Recognition of Distant Conversational Speech

ICASSP 2019accepted

This paper investigates the use of subband temporal envelope (STE) features and speed perturbation based data augmentation in end-to-end recognition of distant conversational speech in everyday home environments. STE features track energy peaks in perceptual frequency bands which reflect the resonan…

Cited by 0SourceScholar