ICASSP 2018accepted0 citations
Filter-and-Convolve: A Cnn Based Multichannel Complex Concatenation Acoustic Model
Abstract
We propose a convolutional neural network (CNN) based multichannel complex-domain concatenation acoustic model. The proposed model extracts speech-specific information from multichannel noisy speech signals. In addition, we design two CNN templates that have wide applicability and several speaker adaptation methods for the multichannel complex concatenation acoustic model. Even with a simple BeamformIt beamformer and the baseline language model, our method obtains a word error rate (WER) of 5.39% on the CHiME-4 corpus, outperforming the previous best result by 13.06% relatively. Using an MVDR beamformer, our model outperforms the corresponding best system by 9.77% relatively.
BibTeX
@inproceedings{icassp2018_filterandconvolv,
title = {Filter-and-Convolve: A Cnn Based Multichannel Complex Concatenation Acoustic Model},
author = {Peidong Wang and DeLiang Wang},
booktitle = {ICASSP 2018},
year = {2018}
}