An Audio Scene Classification Framework with Embedded Filters and a DCT-based Temporal Module
Hangting Chen, Pengyuan Zhang, Yonghong Yan
Abstract
Deep convolutional neural network (DCNN) has recently improved the performance of acoustic scene classification. However, the input features of the network are usually based on predefined hand-tailored filters, which may not apply to the specific tasks. To overcome this, we propose a hybrid framework that jointly trains the front-end filters and the back-end DCNN. Also, a novel temporal module based on the discrete cosine transform (DCT) is inserted after the high-level feature map of the network, thus enabling us to utilize time information without a reduction of training samples. Our single system, composed of the fine-tuned wavelet front-end and the DCNN back-end, with the integrated DCT-based temporal module, has achieved an accuracy of 79.20% in the evaluation set in DCASE17, gaining around 3% and 8% accuracy improvement compared with scalogram-DCNN and FBank-DCNN systems, respectively.
BibTeX
@inproceedings{icassp2019_anaudiosceneclas,
title = {An Audio Scene Classification Framework with Embedded Filters and a DCT-based Temporal Module},
author = {Hangting Chen and Pengyuan Zhang and Yonghong Yan},
booktitle = {ICASSP 2019},
year = {2019}
}