← Search

Yongkang Yin

2 accepted papers

2025

Audio-Faces Intra-Frame Alignment with Graph Attention Networks for Active Speaker Detection

ICASSP 2025accepted

Audio-Visual Active Speaker Detection(ASD) is the task of identifying, at any given moment, who is actively speaking in a multi-person scene by using audio and visual cues. Current main stream ASD methods separately encode audio and facial features, then adopt post-feature fusion approach where the…

Cited by 0SourceScholar