ICASSP 2025accepted0 citations

Leveraging Boolean Directivity Embedding for Binaural Target Speaker Extraction

Yichi Wang, Jie Zhang, Chengqian Jiang, Weitai Zhang, Zhongyi Ye, Lirong Dai

Abstract

Direction-based target speaker extraction (TSE) attracts a constant attention due to the convenience of direction acquisition over assistive video or enrollment audio. The direction clue heavily affects the TSE performance, which might be more seriously in the case of binaural setups due to the small-sized and irregular microphone array. In this paper, we propose Boolean Directivity Embedding (BDE) as a new direction feature in order to precisely lock onto the target speaker independent on microphone array configurations for binaural TSE (BiTSE). We design an encoder that accurately aligns the BDE with the mixed audio signals for feature fusion. Considering that the Boolean representation may contain insufficient spatial and temporal information, we enhance the BDE by incorporating the previously-proposed spatiotemporal features, showing a compatibility and stronger capacity for BiTSE. The proposed BiTSE model, which is based on the narrow-band Conformer as the backbone, can adapt to the cases of target speaker switching and moving by simply modifying the frame-wise BDE. Experimental results demonstrate the efficacy of our method in both stationary and dynamic scenarios. The proposed method is open-sourced in https://github.com/ichi131/Direction-based-BiTSE.

BibTeX
@inproceedings{icassp2025_leveragingboolea,
  title = {Leveraging Boolean Directivity Embedding for Binaural Target Speaker Extraction},
  author = {Yichi Wang and Jie Zhang and Chengqian Jiang and Weitai Zhang and Zhongyi Ye and Lirong Dai},
  booktitle = {ICASSP 2025},
  year = {2025}
}