LOCSELECT: Target Speaker Localization with an Auditory Selective Hearing Mechanism
Yu Chen, Xinyuan Qian, Zexu Pan, Kainan Chen, Haizhou Li
Abstract
The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of speakers. In this paper, we present a target speaker localization algorithm with a selective hearing mechanism. Given a reference speech of the target speaker, we first produce a speaker-dependent spectrogram mask to eliminate interfering speakers’ speech. Subsequently, a Long-Short-Term Memory (LSTM) network is employed to extract the target speaker’s location from the filtered spectrogram. Experiments validate the superiority of our proposed method over existing algorithms for different scale-invariant signal-to-noise ratios (SNR) conditions. Specifically, at SNR = -10 dB, our proposed network LocSelect achieves a mean absolute error (MAE) of 3.55° and an accuracy (ACC) of 87.40%.
BibTeX
@inproceedings{icassp2024_locselecttargets,
title = {LOCSELECT: Target Speaker Localization with an Auditory Selective Hearing Mechanism},
author = {Yu Chen and Xinyuan Qian and Zexu Pan and Kainan Chen and Haizhou Li},
booktitle = {ICASSP 2024},
year = {2024}
}