ICASSP 2018accepted0 citations

On Spatial Features for Supervised Speech Separation and its Application to Beamforming and Robust ASR

Zhong-Qiu Wang, DeLiang Wang

Abstract

This study integrates complementary spectral and spatial information to elevate deep learning based time-frequency masking and acoustic beamforming. Coherence and directional features are designed as additional input features for deep neural network training to remove diffuse noise and other directional interferences pervasive in real-world recordings. The diffuse and directional features are designed to be relatively invariant to the underlying target direction, number of microphones and microphone geometry. The estimated masks are then utilized to compute steering vectors and spatial covariance matrices for beamforming and robust ASR. Experiments on the CHiME-4 dataset demonstrate the effectiveness of the proposed approach.

BibTeX
@inproceedings{icassp2018_onspatialfeature,
  title = {On Spatial Features for Supervised Speech Separation and its Application to Beamforming and Robust ASR},
  author = {Zhong-Qiu Wang and DeLiang Wang},
  booktitle = {ICASSP 2018},
  year = {2018}
}