ICASSP 2023accepted0 citations

Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations

Vasudha Kowtha, Miquel Espi Marques, Jonathan Huang, Yichi Zhang, Carlos Avendaño

Abstract

This work investigates pretrained audio representations for few shot Sound Event Detection. We specifically address the task of few shot detection of novel acoustic sequences, or sound events with semantically meaningful temporal structure, without assuming access to non-target audio. We develop procedures for pretraining suitable representations, and methods which transfer them to our few shot learning scenario. Our experiments evaluate the general purpose utility of our pretrained representations on AudioSet, and the utility of proposed few shot methods via tasks constructed from real-world acoustic sequences. Our pretrained embeddings are suitable to the proposed task, and enable multiple aspects of our few shot framework.

BibTeX
@inproceedings{icassp2023_learningtodetect,
  title = {Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations},
  author = {Vasudha Kowtha and Miquel Espi Marques and Jonathan Huang and Yichi Zhang and Carlos Avendaño},
  booktitle = {ICASSP 2023},
  year = {2023}
}