ICASSP 2022accepted0 citations
Channel-Wise AV-Fusion Attention for Multi-Channel Audio-Visual Speech Recognition
Gaopeng Xu, Song Yang, Wei Li, Song Wang, Guo Wei, Junfeng Yuan, Jie Gao
Abstract
In this paper, we present our work for automatic speech recognition (ASR) in the Multimodal Information Based Speech Processing (MISP) Challenge 2021. We proposed a combination of the guided source separation-based (GSS) speech enhancement technique and a novel Channel-wise Av-fusion encoder (CAE) based acoustic model and found that a kindly combination of these techniques provided essential accuracy improvements. Our ASR system reduces the Chinese Character Error Rate (CCER) by 37.67% absolute compared to the baseline in track 2, achieving first place in the evaluation period with the CCER of 25.07%.
BibTeX
@inproceedings{icassp2022_channelwiseavfus,
title = {Channel-Wise AV-Fusion Attention for Multi-Channel Audio-Visual Speech Recognition},
author = {Gaopeng Xu and Song Yang and Wei Li and Song Wang and Guo Wei and Junfeng Yuan and Jie Gao},
booktitle = {ICASSP 2022},
year = {2022}
}