← Search

Yakun Zhang

3 accepted papers

2026

PURIFICATION BEFORE FUSION: TOWARD MASK-FREE SPEECH ENHANCEMENT FOR ROBUST AUDIO-VISUAL SPEECH RECOGNITION

ICASSP 2026poster

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse interference into the feature fusion process. To mitigate this, rece…

Cited by 0SourcePDFScholar
2024

Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization

COLING 2024main

Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip reading in cross-speaker scenarios where the speaker identity…

Cited by 1SourcePDFScholar
2023

Grounded Entity-Landmark Adaptive Pre-Training for Vision-and-Language Navigation

ICCV 2023oral

Cross-modal alignment is one key challenge for Vision-and-Language Navigation (VLN). Most existing studies concentrate on mapping the global instruction or single sub-instruction to the corresponding trajectory. However, another critical problem of achieving fine-grained alignment at the entity leve…

Cited by 21PDFcodeScholar