← Search

Simon Welker

9 accepted papers

2025

FlowDec: A flow-based full-band general audio codec with high perceptual quality

ICLR 2025poster

We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared to the prior work ScoreDec which is based on score matching, we generalize from…

2024

A Flexible Online Framework for Projection-Based Stft Phase Retrieval

ICASSP 2024accepted

Several recent contributions in the field of iterative STFT phase retrieval have demonstrated that the performance of the classical Griffin-Lim method can be considerably improved upon. By using the same projection operators as Griffin-Lim, but combining them in innovative ways, these approaches ach…

Cited by 0SourceScholar
2024

EMOCONV-Diff: Diffusion-Based Speech Emotion Conversion for Non-Parallel and in-the-Wild Data

ICASSP 2024accepted

Speech emotion conversion is the task of converting the expressed emotion of a spoken utterance to a target emotion while preserving the lexical content and speaker identity. While most existing works in speech emotion conversion rely on acted-out datasets and parallel data samples, in this work we…

Cited by 0SourceScholar
2024

Live Iterative Ptychography with Projection-Based Algorithms

ICASSP 2024accepted

In this work, we demonstrate that the ptychographic phase problem can be solved in a live fashion during scanning, while data is still being collected. We propose a generally applicable modification of the widespread projection-based algorithms such as Error Reduction (ER) and Difference Map (DM). T…

Cited by 0SourceScholar
2023

Analysing Diffusion-based Generative Approaches Versus Discriminative Approaches for Speech Restoration

ICASSP 2023accepted

Diffusion-based generative models have had a high impact on the computer vision and speech processing communities these past years. Besides data generation tasks, they have also been employed for data restoration tasks like speech enhancement and dereverberation. While discriminative models have tra…

Cited by 0SourceScholar
2023

Speech Signal Improvement Using Causal Generative Diffusion Models

ICASSP 2023accepted

In this paper, we present a causal speech signal improvement system that is designed to handle different types of distortions. The method is based on a generative diffusion model which has been shown to work well in scenarios with missing data and non-linear corruptions. To guarantee causal processi…

Cited by 0SourceScholar