ICASSP 2025accepted0 citations

EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible Candidates

Yongqian Li, Xin Zhou, Zheng He, Wei Yu, Yong Luo

Abstract

Active Speaker Detection (ASD) aims to determine whether each candidate in a video frame is speaking. The egocentric dataset Ego4D introduces unique challenges for this task, such as dynamic shooting angles that cause candidates to frequently leave the sight, leading to temporal discontinuities. Additionally, Ego4D poses a novel task: detecting the speaking activities of the camera wearer, who never appears in the field of view. Existing methods treat these two tasks separately, and treat candidates out of sight as noise. In contrast, we propose EgoNet, a framework that uniformly models all candidates, including those not visible. By capturing interactions among all candidates and modeling broader temporal context, EgoNet reduces uncertainty and improves performance in egocentric active speaker detection.

BibTeX
@inproceedings{icassp2025_egonetanunifiede,
  title = {EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible Candidates},
  author = {Yongqian Li and Xin Zhou and Zheng He and Wei Yu and Yong Luo},
  booktitle = {ICASSP 2025},
  year = {2025}
}