← Search

Khanh Le

2 accepted papers

2025

ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription

ICASSP 2025accepted

Deploying ASR models at an industrial scale poses significant challenges in hardware resource management, especially for long-form transcription tasks where audio may last for hours. Large Conformer models, despite their capabilities, are limited to processing only 15 minutes of audio on an 80GB GPU…

Cited by 0SourceScholar
2025

SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition

ICASSP 2025accepted

RNN-Transducer (RNN-T) is a widely adopted architecture in speech recognition, integrating acoustic and language modeling in an end-to-end framework. However, the RNN-T predictor tends to over-rely on consecutive word dependencies in training data, leading to high deletion error rates, particularly…

Cited by 0SourceScholar