Single-channel Speech Extraction Using Speaker Inventory and Attention Network
Xiong Xiao, Zhuo Chen, Takuya Yoshioka, Hakan Erdogan, Changliang Liu, Dimitrios Dimitriadis, Jasha Droppo, Yifan Gong
Abstract
Neural network-based speech separation has received a surge of interest in recent years. Previously proposed methods either are speaker independent or extract a target speaker's voice by using his or her voice snippet. In applications such as home devices or office meeting transcriptions, a possible speaker list is available, which can be leveraged for speech separation. This paper proposes a novel speech extraction method that utilizes an inventory of voice snippets of possible interfering speakers, or speaker enrollment data, in addition to that of the target speaker. Furthermore, an attention-based network architecture is proposed to form time-varying masks for both the target and other speakers during the separation process. This architecture does not reduce the enrollment audio of each speaker into a single vector, thereby allowing each short time frame of the input mixture signal to be aligned and accurately compared with the enrollment signals. We evaluate the proposed system on a speaker extraction task derived from the Libri corpus and show the effectiveness of the method.
BibTeX
@inproceedings{icassp2019_singlechannelspe,
title = {Single-channel Speech Extraction Using Speaker Inventory and Attention Network},
author = {Xiong Xiao and Zhuo Chen and Takuya Yoshioka and Hakan Erdogan and Changliang Liu and Dimitrios Dimitriadis and Jasha Droppo and Yifan Gong},
booktitle = {ICASSP 2019},
year = {2019}
}