The segregation of spatialised speech in interference by optimal mapping of diverse cues
Abstract
We describe optimal cue mapping (OCM), a potentially eal-time binaural signal processing method for segregating sound source in the presence of multiple interfering 3D ound sources. Spatial cues are extracted from a multisource inaural mixture and used to train artificial neural etworks (ANNs) to estimate the spectral energy fraction of wanted speech source in the mixture. Once trained, the NN outputs form a spectral ratio mask which is applied rame-by-frame to the mixture to approximate the agnitude spectrum of the wanted speech. The speech ntelligibility performance of the OCM algorithm for nechoic sound sources is evaluated on previously unseen peech mixtures using the STOI automated measures, and ompared with an established reference method. The ptimized integration of multiple cues offers clear erformance benefits and the ability to quantify the relative mportance of each cue will facilitate computationally fficient implementations.
BibTeX
@inproceedings{icassp2015_thesegregationof,
title = {The segregation of spatialised speech in interference by optimal mapping of diverse cues},
author = {Jingbo Gao and Anthony I. Tew},
booktitle = {ICASSP 2015},
year = {2015}
}