Raw Audio Deep Learning Filter Banks for Acoustic Scene Classification
Abstract
An acoustic scene classification (ASC) framework based on filter banks and raw audio signals is proposed. The time-domain waveform captured by a microphone is initially processed by a filter bank to produce raw audio subbands, which are then processed by a deep neural network (DNN) to estimate the scene class. The network model consists of parallel convolutional neural network (CNN) branches and utilizes late fusion for output class prediction. Each CNN pipeline is designed to handle both the time-domain waveform and the raw audio subbands. The filter bank enhances the frequency components that are crucial for scene recognition, thereby improving classification accuracy. Experiments conducted using the DCASE 2019 dataset show a significant performance improvement compared to models that use raw audio or log-mel spectrogram inputs.
BibTeX
@inproceedings{icassp2025_rawaudiodeeplear,
title = {Raw Audio Deep Learning Filter Banks for Acoustic Scene Classification},
author = {Daniele Salvati},
booktitle = {ICASSP 2025},
year = {2025}
}