ICASSP 2023accepted0 citations

WL-MSR: Watch and Listen for Multimodal Subtitle Recognition

Jiawei Liu, Hao Wang, Weining Wang, Xingjian He, Jing Liu

Abstract

Video subtitles could be defined as the combination of visualized subtitles in frames and textual content recognized from speech, which play a significant role in video understanding for both humans and machines. In this paper, we propose a novel Watch and Listen for Multimodal Subtitle Recognition (WL-MSR) framework to obtain comprehensive video subtitles, by fusing the information provided by Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) models. Specifically, we build a Transformer model with mask and crop strategies and multi-level identity embeddings to aggregate both the textual results and features of the two modalities. To pre-filter out the noise items in OCR results before fusion, we adopt an OCR filter based on ASR results and confidence scores of OCR. By combining these techniques, our solution wins the 2nd place in Multimodal Subtitle Recognition Challenge on ICPR2022.

BibTeX
@inproceedings{icassp2023_wlmsrwatchandlis,
  title = {WL-MSR: Watch and Listen for Multimodal Subtitle Recognition},
  author = {Jiawei Liu and Hao Wang and Weining Wang and Xingjian He and Jing Liu},
  booktitle = {ICASSP 2023},
  year = {2023}
}
WL-MSR: Watch and Listen for Multimodal Subtitle Recognition · ICASSP 2023