2023
Practice of the Conformer Enhanced Audio-Visual Hubert on Mandarin and English
ICASSP 2023accepted
Considering the bimodal nature of human speech perception, lips, and teeth movement has a pivotal role in automatic speech recognition. Benefiting from the correlated and noise-invariant visual information, audio-visual recognition systems enhance robustness in multiple scenarios. In previous work,…