A deep scattering spectrum - Deep Siamese network pipeline for unsupervised acoustic modeling
Neil Zeghidour, Gabriel Synnaeve, Maarten Versteegh, Emmanuel Dupoux
Abstract
Recent work has explored deep architectures for learning acoustic features in an unsupervised or weakly-supervised way for phone recognition. Here we investigate the role of the input features, and in particular we test whether standard mel-scaled filterbanks could be replaced by inherently richer representations, such as derived from an analytic scattering spectrum. We use a Siamese network using lexical side information similar to a well-performing architecture used in the Zero Resource Speech Challenge (2015), and show a substantial improvement when the filterbanks are replaced by scattering features, even though these features yield similar performance when tested without training. This shows that unsupervised and weakly-supervised architectures can benefit from richer features than the traditional ones.
BibTeX
@inproceedings{icassp2016_adeepscatterings,
title = {A deep scattering spectrum - Deep Siamese network pipeline for unsupervised acoustic modeling},
author = {Neil Zeghidour and Gabriel Synnaeve and Maarten Versteegh and Emmanuel Dupoux},
booktitle = {ICASSP 2016},
year = {2016}
}