← Search

Jean-Marie Lemercier

8 accepted papers

2026

Scaling Beyond Masked Diffusion Language Models

ICML 2026poster

Diffusion language models are a promising alternative to autoregressive models due to their potential for faster generation. Among discrete diffusion approaches, Masked diffusion currently dominates, largely driven by strong perplexity on language modeling benchmarks. In this work, we present the fi…

Cited by 0SourceScholar
2025

HRTF Estimation using a Score-based Prior

ICASSP 2025accepted

We present a head-related transfer function (HRTF) estimation method which relies on a data-driven prior given by a score-based diffusion model. The HRTF is estimated in reverberant environments using natural excitation signals, e.g. human speech. The impulse response of the room is estimated along…

Cited by 0SourceScholar
2024

An Independence-promoting Loss for Music Generation with Language Models

ICML 2024poster

Music generation schemes using language modeling rely on a vocabulary of audio tokens, generally provided as codes in a discrete latent space learnt by an auto-encoder. Multi-stage quantizers are often employed to produce these tokens, therefore the decoding strategy used for token prediction must b…

Cited by 3SourcePDFScholar
2024

Single and Few-Step Diffusion for Generative Speech Enhancement

ICASSP 2024accepted

Diffusion models have shown promising results in single-channel speech enhancement, using a task-adapted diffusion process for the conditional generation of clean speech given a noisy mixture. However, at test time, the neural network used for score estimation is called multiple times to solve the i…

Cited by 0SourceScholar
2023

Analysing Diffusion-based Generative Approaches Versus Discriminative Approaches for Speech Restoration

ICASSP 2023accepted

Diffusion-based generative models have had a high impact on the computer vision and speech processing communities these past years. Besides data generation tasks, they have also been employed for data restoration tasks like speech enhancement and dereverberation. While discriminative models have tra…

Cited by 0SourceScholar
2023

Speech Signal Improvement Using Causal Generative Diffusion Models

ICASSP 2023accepted

In this paper, we present a causal speech signal improvement system that is designed to handle different types of distortions. The method is based on a generative diffusion model which has been shown to work well in scenarios with missing data and non-linear corruptions. To guarantee causal processi…

Cited by 0SourceScholar
2022

Customizable End-To-End Optimization Of Online Neural Network-Supported Dereverberation For Hearing Devices

ICASSP 2022accepted

This work focuses on online dereverberation for hearing devices using the weighted prediction error (WPE) algorithm. WPE filtering requires an estimate of the target speech power spectral density (PSD). Recently deep neural networks (DNNs) have been used for this task. However, these approaches opti…

Cited by 0SourceScholar
2021

Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural Network

NeurIPS 2021poster

Recent advances in the design of neural network architectures, in particular those specialized in modeling sequences, have provided significant improvements in speech separation performance. In this work, we propose to use a bio-inspired architecture called Fully Recurrent Convolutional Neural Netwo…