ICASSP 2025accepted0 citations

Scaling Bioacoustic Signal Pre-training with Million Samples Via Mask-Modeling

Xuyao Deng, Tianjiao Wan, Kele Xu, Tian Gao, Peng Qiao, Dawei Feng, Yong Dou

Abstract

Deep learning-based bioacoustic audio analysis holds immense potential across various applications. However, existing studies in bioacoustics often focus on a limited number of species, potentially hindering the transferability of models across different species. Furthermore, the manual annotation of bioacoustic data is both costly and labor-intensive. To address these challenges, self-supervised learning on large-scale bioacoustic audio data presents a promising solution. In this paper, we introduce GPM-BT (General Pre-training Model for Bioacoustic Tasks), a self-supervised, Transformer-based model pre-trained on approximately 1.2 million unannotated bioacoustic audio samples. We evaluate the scalability and effectiveness of this pre-training approach through comprehensive experiments across a broad range of classification and detection tasks. Our results demonstrate that pre-training on large-scale bioacoustic data significantly enhances model performance, improving both generalization and robustness. Notably, GPM-BT achieves state-of-the-art performance on the BEANS benchmark and secures first place in the Few-shot Bioacoustic Event Detection task at the IEEE DCASE 2024 Challenge<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>. To further advance research in bioacoustics, we have open-sourced our models and code<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>.

BibTeX
@inproceedings{icassp2025_scalingbioacoust,
  title = {Scaling Bioacoustic Signal Pre-training with Million Samples Via Mask-Modeling},
  author = {Xuyao Deng and Tianjiao Wan and Kele Xu and Tian Gao and Peng Qiao and Dawei Feng and Yong Dou},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Scaling Bioacoustic Signal Pre-training with Million Samples Via Mask-Modeling · ICASSP 2025