ICASSP 2024accepted0 citations

Cross Modal Training for ASR Error Correction with Contrastive Learning

Jin Jiang, Xiaojun Wan, Wei Peng, Rongjun Li, Jingyuan Yang, Yanquan Zhou

Abstract

ASR Error Correction (AEC) aims to post-process the output of ASR systems and further reduce the word error rate. In this paper, we propose a cross-modal training framework with contrastive learning on the AEC task. This framework enables a shared encoder-decoder model to learn text, pinyin (phoneme <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> ) and audio information simultaneously, which is trained by three subtasks: text correction, pinyin to text and ASR. On this basis, we introduce contrastive learning loss to shrink the distance between the three modalities and construct a unified representation. Experiments <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> on four AEC datasets show that our method effectively corrects a large number of ASR errors to state-of-the-art levels.

BibTeX
@inproceedings{icassp2024_crossmodaltraini,
  title = {Cross Modal Training for ASR Error Correction with Contrastive Learning},
  author = {Jin Jiang and Xiaojun Wan and Wei Peng and Rongjun Li and Jingyuan Yang and Yanquan Zhou},
  booktitle = {ICASSP 2024},
  year = {2024}
}