← Search

Florian Schmid

2 accepted papers

2025

Effective Pre-Training of Audio Transformers for Sound Event Detection

ICASSP 2025accepted

We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously designed training routine on AudioSet frame-level annotations. This includes a balanced sampler, aggressive data augmentation,…

Cited by 0SourceScholar
2023

Efficient Large-Scale Audio Tagging Via Transformer-to-CNN Knowledge Distillation

ICASSP 2023accepted

Audio Spectrogram Transformer models rule the field of Audio Tagging, outrunning previously dominating Convolutional Neural Networks (CNNs). Their superiority is based on the ability to scale up and exploit large-scale datasets such as AudioSet. However, Transformers are demanding in terms of model…

Cited by 0SourceScholar