← Search

Sebastian Braun

19 accepted papers

2024

Adapting Frechet Audio Distance for Generative Music Evaluation

ICASSP 2024accepted

The growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correlation with perceptual quality is understudied. We show that FAD performance may b…

Cited by 0SourceScholar
2023

Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant Environments

ICASSP 2023accepted

Real-time single-channel speech separation aims to unmix an audio stream captured from a single microphone that contains multiple people talking at once, environmental noise, and reverberation into multiple de-reverberated and noise-free speech tracks, each track containing only one talker. While la…

Cited by 0SourceScholar
2022

ICASSP 2022 Acoustic Echo Cancellation Challenge

ICASSP 2022accepted

The ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC challenge and it is enhanced by including mobile scenarios,…

Cited by 83SourceScholar
2022

Icassp 2022 Deep Noise Suppression Challenge

ICASSP 2022accepted

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-sourc…

Cited by 0SourceScholar
2022

Unsupervised Speech Enhancement with Speech Recognition Embedding and Disentanglement Losses

ICASSP 2022accepted

Speech enhancement has recently achieved great success with various deep learning methods. However, most conventional speech enhancement systems are trained with supervised methods that impose two significant challenges. First, a majority of training datasets for speech enhancement systems are synth…

Cited by 0SourceScholar
2021

DBnet: Doa-Driven Beamforming Network for end-to-end Reverberant Sound Source Separation

ICASSP 2021accepted

Many deep learning techniques are available to perform source separation and reduce background noise. However, designing an end-to-end multi-channel source separation method using deep learning and conventional acoustic signal processing techniques still remains challenging. In this paper we propose…

Cited by 0SourceScholar
2021

ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and Results

ICASSP 2021accepted

The ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performa…

Cited by 0SourceScholar
2021

ICASSP 2021 Deep Noise Suppression Challenge

ICASSP 2021accepted

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020 where we open-sourced training and test datasets for researchers to tr…

Cited by 0SourceScholar
2021

Towards Efficient Models for Real-Time Deep Noise Suppression

ICASSP 2021accepted

With recent research advancements, deep learning models are be-coming attractive and powerful choices for speech enhancement in real-time applications. While state-of-the-art models can achieve outstanding results in terms of speech quality and background noise reduction, the main challenge is to ob…

Cited by 0SourceScholar
2020

Joint Beamforming and Reverberation Cancellation Using a Constrained Kalman Filter With Multichannel Linear Prediction

ICASSP 2020accepted

The performance of speech processing systems degrades significantly in far-field scenarios where the distance between the user and microphones increases, leading to low signal-to-noise and signal-to-reverberation ratios. To address this challenge, combining the denoising and dereverberation techniqu…

Cited by 0SourceScholar
2020

Predicting Word Error Rate for Reverberant Speech

ICASSP 2020accepted

Reverberation negatively impacts the performance of automatic speech recognition (ASR). Prior work on quantifying the effect of reverberation has shown that clarity (C50), a parameter that can be estimated from the acoustic impulse response, is correlated with ASR performance. In this paper we propo…

Cited by 0SourceScholar
2020

Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech Enhancement

ICASSP 2020accepted

This paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channel speech enhancement. Specifically, we focus on a RNN that enhances short-time speech spectra on a single-frame-in, sing…

Cited by 0SourceScholar
2018

Dual-Channel Modulation Energy Metric for Direct-to-Reverberation Ratio Estimation

ICASSP 2018accepted

Non-intrusive estimators for acoustic parameters like the direct-to-reverberation ratio (DRR) are useful tools but still perform weakly as shown in the acoustic characterization of environments (ACE) challenge. In this paper, we develop a novel dual-channel metric based on the modulation energy doma…

Cited by 0SourceScholar
2015

Residual noise control using a parametric multichannel Wiener filter

ICASSP 2015accepted

Multichannel noise reduction techniques are commonly used in speech communication applications. In these applications, it is often desired to maintain a residual amount of background noise to avoid perceptually unpleasant artifacts, such as musical tones or time periods of complete silence. Noise re…

Cited by 0SourceScholar