Adaptive extraction of repeating non-negative temporal patterns for single-channel speech enhancement
Yinan Li, Xiongwei Zhang, Meng Sun, Gang Min, Jibin Yang
Abstract
Estimating unknown background noise from single-channel noisy speech is a key yet challenging problem for speech enhancement. Given the fact that the background noises typically have the repeating property and the foreground speech is sparse and time-variant, many literatures decompose the noisy spectrogram directly in an unsupervised fashion when there is no isolated training example of the target speaker or particular noise types beforehand. However, recently proposed methods suffer from un-interpretable decomposed patterns, neglecting the temporal structure of the background noise or being constrained by the pre-fixed parameters. To settle these issues, we propose a novel method based on autocorrelation technique and convolutive non-negative matrix factorization. The proposed method can adaptively estimate the underlying non-negative repeating temporal patterns from noisy speech and identify the clean speech spectrogram simultaneously. Experiments on NOIZEUS dataset mixed with various real-world background noises showed that the proposed method performs better than some state-of-the-art methods.
BibTeX
@inproceedings{icassp2016_adaptiveextracti,
title = {Adaptive extraction of repeating non-negative temporal patterns for single-channel speech enhancement},
author = {Yinan Li and Xiongwei Zhang and Meng Sun and Gang Min and Jibin Yang},
booktitle = {ICASSP 2016},
year = {2016}
}