← Search

Janek Ebbers

7 accepted papers

2025

Keeping the Balance: Anomaly Score Calculation for Domain Generalization

ICASSP 2025accepted

Emitted sounds may drastically change when using different microphones, when properties of the sound sources change, or when recording in different acoustic environments. Ideally, anomalous sound detection (ASD) systems should be able to generalize well to unseen target domains by only providing a f…

Cited by 0SourceScholar
2025

Leveraging Audio-Only Data for Text-Queried Target Sound Extraction

ICASSP 2025accepted

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text queries, the limited number of available high-quality text-…

Cited by 0SourceScholar
2025

No Class Left Behind: A Closer Look at Class Balancing for Audio Tagging

ICASSP 2025accepted

Large-scale audio tagging datasets like AudioSet usually suffer from severe class imbalance comprising many audio examples for common sound classes but only few examples of rare sound classes. The latter, however, may yet be equally or even more important to recognize. Therefore, it is common practi…

Cited by 0SourceScholar
2025

Task-Aware Unified Source Separation

ICASSP 2025accepted

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or cinematic audio source separation (CASS) with a single model. These models are trained on large-scale data including spe…

Cited by 0SourceScholar
2025

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing

CVPR 2025poster

Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in both modalities concurrently). Moreover, the prohibitive cost o…

2021

Contrastive Predictive Coding Supported Factorized Variational Autoencoder For Unsupervised Learning Of Disentangled Speech Representations

ICASSP 2021accepted

In this work we address disentanglement of style and content in speech signals. We propose a fully convolutional variational autoencoder employing two encoders: a content encoder and a style encoder. To foster disentanglement, we propose adversarial contrastive predictive coding. This new disentangl…

Cited by 24SourceScholar