← Search

Julius Richter

9 accepted papers

2026

ARTIFREE: DETECTING AND REDUCING GENERATIVE ARTIFACTS IN DIFFUSION-BASED SPEECH ENHANCEMENT

ICASSP 2026poster

Diffusion-based speech enhancement (SE) achieves natural-sounding speech and strong generalization, yet suffers from key limitations like generative artifacts and high inference latency. In this work, we systematically study artifact prediction and reduction in diffusion-based SE. We show that varia…

Cited by 0SourcePDFScholar
2026

Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture

ICASSP 2026poster

We study diffusion-based speech enhancement using a Schrodinger bridge formulation and extend the EDM2 framework to this setting. We employ time-dependent preconditioning of network inputs and outputs to stabilize training and explore two skip-connection configurations that allow the network to pred…

Cited by 0SourcePDFScholar
2026

Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning

CVPR 2026

We introduce Perception Encoder-Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Building on PE, PE-AV makes several key contributions to extend representations to audio, and natively support joint embeddings across audio-video,

Cited by 0SourcecodeScholar
2025

Investigating Training Objectives for Generative Speech Enhancement

ICASSP 2025accepted

Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these framewor…

Cited by 0SourceScholar
2024

Single and Few-Step Diffusion for Generative Speech Enhancement

ICASSP 2024accepted

Diffusion models have shown promising results in single-channel speech enhancement, using a task-adapted diffusion process for the conditional generation of clean speech given a noisy mixture. However, at test time, the neural network used for score estimation is called multiple times to solve the i…

Cited by 0SourceScholar
2023

Analysing Diffusion-based Generative Approaches Versus Discriminative Approaches for Speech Restoration

ICASSP 2023accepted

Diffusion-based generative models have had a high impact on the computer vision and speech processing communities these past years. Besides data generation tasks, they have also been employed for data restoration tasks like speech enhancement and dereverberation. While discriminative models have tra…

Cited by 0SourceScholar
2023

Speech Signal Improvement Using Causal Generative Diffusion Models

ICASSP 2023accepted

In this paper, we present a causal speech signal improvement system that is designed to handle different types of distortions. The method is based on a generative diffusion model which has been shown to work well in scenarios with missing data and non-linear corruptions. To guarantee causal processi…

Cited by 0SourceScholar
2021

Guided Variational Autoencoder for Speech Enhancement with a Supervised Classifier

ICASSP 2021accepted

Recently, variational autoencoders have been successfully used to learn a probabilistic prior over speech signals, which is then used to perform speech enhancement. However, variational autoencoders are trained on clean speech only, which results in a limited ability of extracting the speech signal…

Cited by 0SourceScholar