ICASSP 2025accepted0 citations

Curriculum Learning aided Audio-Visual Speech Recognition with Arbitrary Speaker Number

Yuxiao Lin, Tao Jin, Xize Cheng, Zhou Zhao, Fei Wu

Abstract

Recently, audio-visual speech recognition has attracted increasing attention. However, most existing works only focused on scenarios with two speakers. In this work, we study the effect of speaker number in AVSR task and propose an end-to-end audio-visual speech recognition framework under a more realistic condition where the speaker number is arbitrary. Specifically, we adopted curriculum learning to train models from easy scenarios to hard ones and introduce a new training strategy named Challenge-based Curriculum Learning (CBCL) that forces the model to focus on hard, challenging data instead of easy ones during training. Further, to avoid scenario bias from unbalanced sampling during curriculum learning, we propose a Speaker-number Aware Mixture-of-Expert (SA-MoE) mechanism to explicitly model the characteristic difference in scenarios with different speaker numbers.

BibTeX
@inproceedings{icassp2025_curriculumlearni,
  title = {Curriculum Learning aided Audio-Visual Speech Recognition with Arbitrary Speaker Number},
  author = {Yuxiao Lin and Tao Jin and Xize Cheng and Zhou Zhao and Fei Wu},
  booktitle = {ICASSP 2025},
  year = {2025}
}