Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system
Jahn Heymann, Lukas Drude, Christoph Böddeker, Patrick Hanebrink, Reinhold Haeb-Umbach
Abstract
This paper presents an end-to-end training approach for a beamformer-supported multi-channel ASR system. A neural network which estimates masks for a statistically optimum beamformer is jointly trained with a network for acoustic modeling. To update its parameters, we propagate the gradients from the acoustic model all the way through feature extraction and the complex valued beamforming operation. Besides avoiding a mismatch between the front-end and the back-end, this approach also eliminates the need for stereo data, i.e., the parallel availability of clean and noisy versions of the signals. Instead, it can be trained with real noisy multi-channel data only. Also, relying on the signal statistics for beamforming, the approach makes no assumptions on the configuration of the microphone array. We further observe a performance gain through joint training in terms of word error rate in an evaluation of the system on the CHiME 4 dataset.
BibTeX
@inproceedings{icassp2017_beamnetendtoendt,
title = {Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system},
author = {Jahn Heymann and Lukas Drude and Christoph Böddeker and Patrick Hanebrink and Reinhold Haeb-Umbach},
booktitle = {ICASSP 2017},
year = {2017}
}