← Search

Wupeng Wang

3 accepted papers

2023

ImagineNet: Target Speaker Extraction with Intermittent Visual Cue Through Embedding Inpainting

ICASSP 2023accepted

The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a pre-recorded utterance or a synchronized lip movement in a video cl…

Cited by 0SourceScholar
2021

Vset: A Multimodal Transformer for Visual Speech Enhancement

ICASSP 2021accepted

The transformer architecture has shown great capability in learning long-term dependency and works well in multiple domains. However, transformer has been less considered in audio-visual speech enhancement (AVSE) research, partly due to the convention that treats speech enhancement as a short-time s…

Cited by 0SourceScholar