ICASSP 2015accepted0 citations

Variational EM for clustering interaural phase cues in MESSL for blind source separation of speech

Zeinab Zohny, Syed Mohsen Naqvi, Jonathon A. Chambers

Abstract

The model-based expectation maximization source separation and localization (MESSL) technique is a probabilistic time-frequency masking algorithm that achieves underdetermined blind source separation of speech sources. Using only two-channel recordings, MESSL clusters spectrogram points based on their interaural spatial cues. Gaussian mixture models (GMMs) are assumed for the interaural cues and their corresponding parameters are determined by maximum likelihood estimation (MLE) via the expectation maximization (EM) framework. However, the presence of singularities and over-fitting are major drawbacks of MLE. In this paper, we investigate variational Bayesian (VB) inference for clustering spectrogram points based particularly on their interaural phase difference (IPD) cues. Variational inference overcomes the difficulties associated with the likelihood optimization and improves the separation especially when the sources are in close proximity. Simulation studies based on speech mixtures formed from the TIMIT database confirm the advantage of the proposed approach in terms of signal to distortion ratio (SDR).

BibTeX
@inproceedings{icassp2015_variationalemfor,
  title = {Variational EM for clustering interaural phase cues in MESSL for blind source separation of speech},
  author = {Zeinab Zohny and Syed Mohsen Naqvi and Jonathon A. Chambers},
  booktitle = {ICASSP 2015},
  year = {2015}
}