2021
Vset: A Multimodal Transformer for Visual Speech Enhancement
ICASSP 2021accepted
The transformer architecture has shown great capability in learning long-term dependency and works well in multiple domains. However, transformer has been less considered in audio-visual speech enhancement (AVSE) research, partly due to the convention that treats speech enhancement as a short-time s…