← Search

Alfred Mertins

15 accepted papers

2023

Improving Automatic Sleep Staging Via Temporal Smoothness Regularization

ICASSP 2023accepted

We propose a regularization method, so-called temporal smoothness regularization, for training deep neural networks for automatic sleep staging in small data settings. In intuition, we constrain the cross-entropy losses of any two adjacent epochs in the sequential input to be as close to each other…

Cited by 0SourceScholar
2022

Polyphonic Audio Event Detection: Multi-Label or Multi-Class Multi-Task Classification Problem?

ICASSP 2022accepted

Polyphonic events are the main error source of audio event detection (AED) systems. In deep-learning context, the most common approach to deal with event overlaps is to treat the AED task as a multi-label classification problem. By doing this, we inherently consider multiple one-vs.-rest classificat…

Cited by 0SourceScholar
2021

Multi-View Audio And Music Classification

ICASSP 2021accepted

We propose in this work a multi-view learning approach for audio and music classification. Considering four typical low-level representations (i.e. different views) commonly used for audio and music recognition tasks, the proposed multi-view network consists of four subnetworks, each handling one in…

Cited by 0SourceScholar
2021

Self-Attention Generative Adversarial Network for Speech Enhancement

ICASSP 2021accepted

Existing generative adversarial networks (GANs) for speech enhancement solely rely on the convolution operation, which may obscure temporal dependencies across the sequence input. To remedy this issue, we propose a self-attention layer adapted from non-local attention, coupled with the convolutional…

Cited by 0SourceScholar
2021

Spherical Harmonic Representation for Dynamic Sound-Field Measurements

ICASSP 2021accepted

Continuously moving microphones produce a high number of spatially dense sound-field samples with low effort in hardware and acquisition time. By interpreting the dynamic procedure as the non-uniform sampling of spatial basis functions, a system of linear equations can be set up. Its solution encode…

Cited by 0SourceScholar
2019

Forked Recurrent Neural Network for Hand Gesture Classification Using Inertial Measurement Data

ICASSP 2019accepted

For many applications of hand gesture recognition, a delay-free, affordable, and mobile system relying on body signals is mandatory. Therefore, we propose an approach for hand gestures classification given signals of inertial measurement units (IMUs) that works with extremely short windows to avoid…

Cited by 0SourceScholar
2019

Unifying Isolated and Overlapping Audio Event Detection with Multi-label Multi-task Convolutional Recurrent Neural Networks

ICASSP 2019accepted

We propose a multi-label multi-task framework based on a convolutional recurrent neural network to unify detection of isolated and overlapping audio events. The framework leverages the power of convolutional recurrent neural network architectures; convolutional layers learn effective features over w…

Cited by 0SourceScholar
2018

Compressive Sampling of Sound Fields Using Moving Microphones

ICASSP 2018accepted

For conventional sampling of sound-fields, the measurement in space by use of stationary microphones is impractical for high audio frequencies. Satisfying the Nyquist-Shannon sampling theorem requires a huge number of sampling points and entails other difficulties, such as the need for exact calibra…

Cited by 0SourceScholar
2018

Weighted and Multi-Task Loss for Rare Audio Event Detection

ICASSP 2018accepted

We present in this paper two loss functions tailored for rare audio event detection in audio streams. The weighted loss is designed to tackle the common issue of imbalanced data in background/foreground classification while the multi-task loss enables the networks to simultaneously model the class d…

Cited by 0SourceScholar
2017

CNN-LTE: A class of 1-X pooling convolutional neural networks on label tree embeddings for audio scene classification

ICASSP 2017accepted

We present in this work an approach for audio scene classification. Firstly, given the label set of the scenes, a label tree is automatically constructed where the labels are grouped into meta-classes. This category taxonomy is then used in the feature extraction step in which an audio scene instanc…

Cited by 0SourceScholar
2017

Measurement of sound fields using moving microphones

ICASSP 2017accepted

The sampling of sound fields involves the measurement of spatially dependent room impulse responses, where the Nyquist-Shannon sampling theorem applies in both the temporal and spatial domains. Therefore, sampling inside a volume of interest requires a huge number of sampling points in space, which…

Cited by 0SourceScholar
2016

Learning compact structural representations for audio events using regressor banks

ICASSP 2016accepted

We introduce a new learned descriptor for audio signals which is efficient for event representation. The entries of the descriptor are produced by evaluating a set of regressors on the input signal. The regressors are class-specific and trained using the random regression forests framework. Given an…

Cited by 0SourceScholar
2015

Joint time- and frequency-domain reshaping of room impulse responses

ICASSP 2015accepted

In listening room compensation, the aim is to compensate for the degradations that are rendered to an audio signal by transmission in a closed room. Due to multiple reflections of the soundwaves, the listener receives a superposition of delayed and attenuated versions of the source signal. A filter…

Cited by 0SourceScholar