Cross Modal Training for ASR Error Correction with Contrastive Learning
Jin Jiang, Xiaojun Wan, Wei Peng, Rongjun Li, Jingyuan Yang, Yanquan Zhou
Abstract
ASR Error Correction (AEC) aims to post-process the output of ASR systems and further reduce the word error rate. In this paper, we propose a cross-modal training framework with contrastive learning on the AEC task. This framework enables a shared encoder-decoder model to learn text, pinyin (phoneme <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> ) and audio information simultaneously, which is trained by three subtasks: text correction, pinyin to text and ASR. On this basis, we introduce contrastive learning loss to shrink the distance between the three modalities and construct a unified representation. Experiments <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> on four AEC datasets show that our method effectively corrects a large number of ASR errors to state-of-the-art levels.
BibTeX
@inproceedings{icassp2024_crossmodaltraini,
title = {Cross Modal Training for ASR Error Correction with Contrastive Learning},
author = {Jin Jiang and Xiaojun Wan and Wei Peng and Rongjun Li and Jingyuan Yang and Yanquan Zhou},
booktitle = {ICASSP 2024},
year = {2024}
}