← Search

Bernd Edler

8 accepted papers

2025

FlowMAC: Conditional Flow Matching for Audio Coding at Low Bit Rates

ICASSP 2025accepted

This paper introduces FlowMAC, a novel neural audio codec for high-quality general audio compression at low bit rates based on conditional flow matching (CFM). FlowMAC jointly learns a mel spectrogram encoder, quantizer and decoder. At inference time the decoder integrates a continuous normalizing f…

Cited by 6SourceScholar
2022

A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain

ICASSP 2022accepted

Frequency domain processing, and in particular the use of Modified Discrete Cosine Transform (MDCT), is the most widespread approach to audio coding. However, at low bitrates, audio quality, especially for speech, degrades drastically due to the lack of available bits to directly code the transform…

Cited by 0SourceScholar
2019

Perceptual Audio Coding with Adaptive Non-uniform Time/frequency Tilings Using Subband Merging and Time Domain Aliasing Reduction

ICASSP 2019accepted

In this paper, we investigate the coding efficiency of perceptual coding using an adaptive non-uniform orthogonal filter-bank based on MDCT analysis/synthesis and time domain aliasing reduction. We compare its performance to a system using a traditional adaptive uniform MDCT filterbank with window s…

Cited by 2SourceScholar
2018

Classification vs. Regression in Supervised Learning for Single Channel Speaker Count Estimation

ICASSP 2018accepted

The task of estimating the maximum number of concurrent speakers from single channel mixtures is important for various audio-based applications, such as blind source separation, speaker diarisation, audio surveillance or auditory scene classification. Building upon powerful machine learning methodol…

Cited by 0SourceScholar
2016

Common fate model for unison source separation

ICASSP 2016accepted

In this paper we present a novel source separation method aiming to overcome the difficulty of modelling non-stationary signals. The method can be applied to mixtures of musical instruments with frequency and/or amplitude modulation, e.g. typically caused by vibrato. It is based on a signal represen…

Cited by 30SourceScholar