ICASSP 2025accepted0 citations

Multi-modal Streaming ASR in Cross-talk Scenario for Smart Glasses

Ya Jiang, Hongbo Lan, Qing Wang, Shutong Niu

Abstract

In the MMCSG task of the CHiME-8 Challenge, achieving real-time speaker-attributed transcriptions with limited multi-modal data presents significant challenges. To cope with the problem, we propose a novel ASR framework that leverages both audio-only and multi-modal inputs in a streaming fashion. For the audio-only modality, analyzing and emulating the characteristics of real audio, we utilize a multi-channel simulation to generate the augmented dataset, which efficiently reduces the model training deviation between real and simulated data. Additionally, we integrate the IMU data with audio data in the network structure, demonstrating that the functional filtered and encoded IMU data can assist audio information in achieving better real-time speech recognition performance with ablation experiments. Notably, our explorations based on the above schemes not only secured first place in the MMCSG sub-track but also represented the first investigation into the effectiveness of leveraging IMU data for this task.

BibTeX
@inproceedings{icassp2025_multimodalstream,
  title = {Multi-modal Streaming ASR in Cross-talk Scenario for Smart Glasses},
  author = {Ya Jiang and Hongbo Lan and Qing Wang and Shutong Niu},
  booktitle = {ICASSP 2025},
  year = {2025}
}