Multichannel Speaker Activity Detection for Meetings
Patrick Meyer, Rolf Jongebloed, Tim Fingscheidt
Abstract
Multichannel recordings of meetings with a (wireless) headset for each person deliver commonly the best audio quality for subsequent analyses. However, still speech portions of other participants can couple into the microphone channel of the associated target speaker. Due to this crosstalk, a speaker activity detection (SAD) is required in order to identify only the speech portions of the target speaker in the related microphone channel. While most solutions are either complex and need a training process, or achieve insufficient results in multi-talk situations, we propose a low complexity method, which can handle both crosstalk and multi-talk situations. We investigate single- and multi-talk in a wide range of different crosstalk levels, and improved the detection accuracy towards a standardized voice activity detection overall by 12.89 % absolute, whereas a state-of-the-art multichannel SAD was exceeded even by 13.76 % absolute.
BibTeX
@inproceedings{icassp2018_multichannelspea,
title = {Multichannel Speaker Activity Detection for Meetings},
author = {Patrick Meyer and Rolf Jongebloed and Tim Fingscheidt},
booktitle = {ICASSP 2018},
year = {2018}
}