← Search

Wolfgang Mack

5 accepted papers

2026

Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions

ICASSP 2026oral

Objective speech-quality metrics are widely used to assess codec performance. However, for neural codecs, it is often unclear which metrics provide reliable quality estimates. To address this, we evaluated 45 objective metrics by correlating their scores with subjective listening scores for clean sp…

Cited by 0SourcePDFScholar
2023

Multi-Microphone Speaker Separation by Spatial Regions

ICASSP 2023accepted

We consider the task of region-based source separation of reverberant multi-microphone recordings. We assume pre-defined spatial regions with a single active source per region. The objective is to estimate the signals from the individual spatial regions as captured by a reference microphone while re…

Cited by 0SourceScholar
2021

Efficient Training Data Generation for Phase-Based DOA Estimation

ICASSP 2021accepted

Deep learning (DL) based direction of arrival (DOA) estimation is an active research topic and currently represents the state-of-the-art. Usually, DL-based DOA estimators are trained with recorded data or computationally expensive generated data. Both data types require significant storage and exces…

Cited by 0SourceScholar
2020

Data-Driven Wind Speed Estimation Using Multiple Microphones

ICASSP 2020accepted

A deep neural network (DNN) based approach for estimating the speed of airflows using closely-spaced microphones is proposed. The spatial characteristics of wind noise measured with a smallaperture array are exploited, i.e., the low-frequency spatial coherence of wind noise signals is used as an inp…

Cited by 2SourceScholar
2020

Signal-Aware Broadband DOA Estimation Using Attention Mechanisms

ICASSP 2020accepted

We refer to direction-of-arrivals (DOAs) estimation of a user-defined subset of directional (desired) sound sources as signal-aware DOA estimation. Source selection, thereby, can be achieved with time-frequency masks to apply attention to TF bins dominated by desired sources. With deep neural networ…

Cited by 0SourceScholar