Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC Loss
Ming Fang, Tao Wei, Kai Guo, Ziyang Zhuang, Yan Shi, Ning Cheng, Shaojun Wang, Jing Xiao
Abstract
Contextual biasing is essential for addressing scenario-specific challenges in End-to-End (E2E) Automatic Speech Recognition (ASR) systems. Prior contextual E2E ASR methods, such as the contextual bias with CPP Network, have utilized bias CTC loss for explicit supervision of bias tasks, However, the alignment between the ASR and bias tasks has been largely neglected. This paper introduces a novel contextual biasing approach that employs the Multi-label Synchronous Output CTC (MCTC) algorithm to enhance the synchronization between ASR and bias task outputs. Furthermore, we propose an enhancement to phrase-level contextual representation. Our proposed method demonstrates significant improvements in Word Error Rate (WER). Specifically, our experiments reveal a 15.0% reduction in WER on the Librispeech-960 dataset compared to the CPPN method, with an impressive 29.0% reduction in WER for context phrases.
BibTeX
@inproceedings{icassp2025_improvingcontext,
title = {Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC Loss},
author = {Ming Fang and Tao Wei and Kai Guo and Ziyang Zhuang and Yan Shi and Ning Cheng and Shaojun Wang and Jing Xiao},
booktitle = {ICASSP 2025},
year = {2025}
}