ICASSP 2021accepted0 citations

Vset: A Multimodal Transformer for Visual Speech Enhancement

Karthik Ramesh, Chao Xing, Wupeng Wang, Dong Wang, Xiao Chen

Abstract

The transformer architecture has shown great capability in learning long-term dependency and works well in multiple domains. However, transformer has been less considered in audio-visual speech enhancement (AVSE) research, partly due to the convention that treats speech enhancement as a short-time signal processing task. In this paper, we challenge this common belief and show that an audio-visual transformer can significantly improve AVSE performance, by learning the long-term dependency of both intra-modality and inter-modality. We test this new transformer-based AVSE model on the GRID and AVSpeech datasets, and show that it beats several state-of-the-art models by a large margin.

BibTeX
@inproceedings{icassp2021_vsetamultimodalt,
  title = {Vset: A Multimodal Transformer for Visual Speech Enhancement},
  author = {Karthik Ramesh and Chao Xing and Wupeng Wang and Dong Wang and Xiao Chen},
  booktitle = {ICASSP 2021},
  year = {2021}
}