← Search

Laurent Girin

15 accepted papers

2025

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

ICASSP 2025accepted

This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key attributes, such as speaker identity, pitch, content, loudness, signal-to-noise rat…

Cited by 0SourceScholar
2023

Speech Modeling with a Hierarchical Transformer Dynamical VAE

ICASSP 2023accepted

The dynamical variational autoencoders (DVAEs) are a family of latent-variable deep generative models that extends the VAE to model a sequence of observed data and a corresponding sequence of latent vectors. In almost all the DVAEs of the literature, the temporal dependencies within each sequence an…

Cited by 0SourceScholar
2022

Repeat after Me: Self-Supervised Learning of Acoustic-to-Articulatory Mapping by Vocal Imitation

ICASSP 2022accepted

We propose a computational model of speech production combining a pre-trained neural articulatory synthesizer able to reproduce complex speech stimuli from a limited set of interpretable articulatory parameters, a DNN-based internal forward model predicting the sensory consequences of articulatory c…

Cited by 0SourceScholar
2020

A Recurrent Variational Autoencoder for Speech Enhancement

ICASSP 2020accepted

This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a nonnegative matrix factorization noise model for speech enhancement. We propose…

Cited by 0SourceScholar
2019

Semi-supervised Multichannel Speech Enhancement with Variational Autoencoders and Non-negative Matrix Factorization

ICASSP 2019accepted

In this paper we address speaker-independent multichannel speech enhancement in unknown noisy environments. Our work is based on a well-established multichannel local Gaussian modeling framework. We propose to use a neural network for modeling the speech spectro-temporal content. The parameters of t…

Cited by 0SourceScholar
2019

Speech Enhancement with Variational Autoencoders and Alpha-stable Distributions

ICASSP 2019accepted

This paper focuses on single-channel semi-supervised speech enhancement. We learn a speaker-independent deep generative speech model using the framework of variational autoencoders. The noise model remains unsupervised because we do not assume prior knowledge of the noisy recording environment. In t…

Cited by 0SourceScholar
2018

Accounting for Room Acoustics in Audio-Visual Multi-Speaker Tracking

ICASSP 2018accepted

Multiple-speaker tracking is a crucial task for many applications. In real-world scenarios, exploiting the complementarity between auditory and visual data enables to track people outside the visual field of view. However, practical methods must be robust to changes in acoustic conditions, e.g. reve…

Cited by 0SourceScholar
2017

An EM algorithm for joint source separation and diarisation of multichannel convolutive speech mixtures

ICASSP 2017accepted

We present a probabilistic model for joint source separation and diarisation of multichannel convolutive speech mixtures. We build upon the framework of local Gaussian model (LGM) with non-negative matrix factorization (NMF). The diarisation is introduced as a temporal labeling of each source in the…

Cited by 0SourceScholar
2017

Audio source separation based on convolutive transfer function and frequency-domain lasso optimization

ICASSP 2017accepted

This paper addresses the problem of under-determined convolutive audio source separation in a semi-oracle configuration where the mixing filters are assumed to be known. We propose a separation procedure based on the convolutive transfer function (CTF), which is a more appropriate model for strongly…

Cited by 0SourceScholar
2016

An inverse-gamma source variance prior with factorized parameterization for audio source separation

ICASSP 2016accepted

In this paper we present a new statistical model for the power spectral density (PSD) of an audio signal and its application to multichannel audio source separation (MASS). The source signal is modeled with the local Gaussian model (LGM) and we propose to model its variance with an inverse-Gamma dis…

Cited by 0SourceScholar
2016

Deep neural networks for automatic detection of screams and shouted speech in subway trains

ICASSP 2016accepted

Deep Neural Networks (DNNs) have recently become a popular technique for regression and classification problems. Their capacity to learn high-order correlations between input and output data proves to be very powerful for automatic speech recognition. In this paper we investigate the use of DNNs for…

Cited by 0SourceScholar
2016

Non-stationary noise power spectral density estimation based on regional statistics

ICASSP 2016accepted

Estimating the noise power spectral density (PSD) is essential for single channel speech enhancement algorithms. In this paper, we propose a noise PSD estimation approach based on regional statistics. The proposed regional statistics consist of four features representing the statistics of the past a…

Cited by 0SourceScholar
2016

Reverberant sound localization with a robot head based on direct-path relative transfer function

IROS 2016poster

This paper addresses the problem of sound-source localization (SSL) with a robot head, which remains a challenge in real-world environments. In particular we are interested in locating speech sources, as they are of high interest for human-robot interaction. The microphone-pair response correspondin…

Cited by 44SourceScholar
2015

Estimation of relative transfer function in the presence of stationary noise based on segmental power spectral density matrix subtraction

ICASSP 2015accepted

This paper addresses the problem of relative transfer function (RTF) estimation in the presence of stationary noise. We propose an RTF identification method based on segmental power spectral density (PSD) matrix subtraction. First multiple channel microphone signals are divided into segments corresp…

Cited by 0SourceScholar