← Search

Simon Leglaive

13 accepted papers

2026

MODELING STRATEGIES FOR SPEECH ENHANCEMENT IN THE LATENT SPACE OF A NEURAL AUDIO CODEC

ICASSP 2026poster

Neural audio codecs (NACs) provide compact latent speech representations in the form of sequences of continuous vectors or discrete tokens. In this work, we investigate how these two types of speech representations compare when used as training targets for supervised speech enhancement. We consider…

Cited by 0SourcePDFScholar
2025

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

ICASSP 2025accepted

This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key attributes, such as speaker identity, pitch, content, loudness, signal-to-noise rat…

Cited by 0SourceScholar
2025

MEGA: Masked Generative Autoencoder for Human Mesh Recovery

CVPR 2025poster

Human Mesh Recovery (HMR) from a single RGB image is a highly ambiguous problem, as an infinite set of 3D interpretations can explain the 2D observation equally well. Nevertheless, most HMR methods overlook this issue and make a single prediction without accounting for this ambiguity. A few approach…

Cited by 1SourcePDFScholar
2024

Towards Improving Speech Emotion Recognition Using Synthetic Data Augmentation from Emotion Conversion

ICASSP 2024accepted

One of the main challenges in speech emotion recognition is the lack of large labelled datasets. The progress in speech synthesis allows us to generate reliable and realistic expressive speech. In this work, we propose using a state-of-the-art end-to-end speech emotion conversion model to generate n…

Cited by 0SourceScholar
2024

VQ-HPS: Human Pose and Shape Estimation in a Vector-Quantized Latent Space

ECCV 2024poster

"Previous works on Human Pose and Shape Estimation (HPSE) from RGB images can be broadly categorized into two main groups: parametric and non-parametric approaches. Parametric techniques leverage a low-dimensional statistical body model for realistic results, whereas recent non-parametric methods ac…

2023

Speech Modeling with a Hierarchical Transformer Dynamical VAE

ICASSP 2023accepted

The dynamical variational autoencoders (DVAEs) are a family of latent-variable deep generative models that extends the VAE to model a sequence of observed data and a corresponding sequence of latent vectors. In almost all the DVAEs of the literature, the temporal dependencies within each sequence an…

Cited by 0SourceScholar
2020

A Recurrent Variational Autoencoder for Speech Enhancement

ICASSP 2020accepted

This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a nonnegative matrix factorization noise model for speech enhancement. We propose…

Cited by 0SourceScholar
2019

Semi-supervised Multichannel Speech Enhancement with Variational Autoencoders and Non-negative Matrix Factorization

ICASSP 2019accepted

In this paper we address speaker-independent multichannel speech enhancement in unknown noisy environments. Our work is based on a well-established multichannel local Gaussian modeling framework. We propose to use a neural network for modeling the speech spectro-temporal content. The parameters of t…

Cited by 0SourceScholar
2019

Speech Enhancement with Variational Autoencoders and Alpha-stable Distributions

ICASSP 2019accepted

This paper focuses on single-channel semi-supervised speech enhancement. We learn a speaker-independent deep generative speech model using the framework of variational autoencoders. The noise model remains unsupervised because we do not assume prior knowledge of the noisy recording environment. In t…

Cited by 0SourceScholar
2018

Alpha-Stable Low-Rank Plus Residual Decomposition for Speech Enhancement

ICASSP 2018accepted

In this study, we propose a novel probabilistic model for separating clean speech signals from noisy mixtures by decomposing the mixture spectra into a structured speech part and a more flexible residual part. The main novelty in our model is that it uses a family of heavy-tailed distributions, so c…

Cited by 0SourceScholar
2017

Alpha-stable multichannel audio source separation

ICASSP 2017accepted

In this paper, we focus on modeling multichannel audio signals in the short-time Fourier transform domain for the purpose of source separation. We propose a probabilistic model based on a class of heavy-tailed distributions, in which the observed mixtures and the latent sources are jointly modeled b…

Cited by 0SourceScholar
2017

Multichannel audio source separation: Variational inference of time-frequency sources from time-domain observations

ICASSP 2017accepted

A great number of methods for multichannel audio source separation are based on probabilistic approaches in which the sources are modeled as latent random variables in a Time-Frequency (TF) domain. For reverberant mixtures, it is common to approximate the time-domain convolutive mixing process as be…

Cited by 0SourceScholar