Multichannel audio source separation: Variational inference of time-frequency sources from time-domain observations
Simon Leglaive, Roland Badeau, Gaël Richard
Abstract
A great number of methods for multichannel audio source separation are based on probabilistic approaches in which the sources are modeled as latent random variables in a Time-Frequency (TF) domain. For reverberant mixtures, it is common to approximate the time-domain convolutive mixing process as being instantaneous in the short-term Fourier transform domain, under a short mixing filters assumption. The TF latent sources are then inferred from the TF mixture observations. In this paper we propose to infer the TF latent sources from the time-domain observations. This approach allows us to exactly model the convolutive mixing process. The inference procedure relies on a variational expectation-maximization algorithm. In significant reverberation conditions, our approach leads to a signal-to-distortion ratio improvement of 5.5 dB compared with the usual TF approximation of the convolutive mixing process.
BibTeX
@inproceedings{icassp2017_multichannelaudi,
title = {Multichannel audio source separation: Variational inference of time-frequency sources from time-domain observations},
author = {Simon Leglaive and Roland Badeau and Gaël Richard},
booktitle = {ICASSP 2017},
year = {2017}
}