ICASSP 2018accepted0 citations

A Casa Approach to Deep Learning Based Speaker-Independent Co-Channel Speech Separation

Yuzhou Liu, DeLiang Wang

Abstract

We address speaker-independent co-channel speech separation from the computational auditory scene analysis (CAS A) perspective. Specifically, we decompose the two-speaker separation task into the stages of simultaneous grouping and sequential grouping. Simultaneous grouping is first performed at the frame level by separating the spectra of two speakers with a permutation-invariantly trained recurrent neural network (RNN). In the second stage, the simultaneously separated spectra at each frame are sequentially grouped into the utterances of the two underlying speakers by a clustering RNN. Overall optimization is then performed to fine tune the two-stage system. The proposed CASA approach takes advantage of permutation invariant training (PIT) and deep clustering (DC), but overcomes their shortcomings. Experiments show that the proposed system improves over the best reported results of PIT and DC.

BibTeX
@inproceedings{icassp2018_acasaapproachtod,
  title = {A Casa Approach to Deep Learning Based Speaker-Independent Co-Channel Speech Separation},
  author = {Yuzhou Liu and DeLiang Wang},
  booktitle = {ICASSP 2018},
  year = {2018}
}