← Search

Gaël RICHARD

46 accepted papers

2026

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

ICML 2026poster

We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference time. Traditional language modeling is an intrinsically ill-posed problem: given a context, multiple ``futures'' may be…

Cited by 0SourceScholar
2026

S-PRESSO: ULTRA LOW BITRATE SOUND EFFECT COMPRESSION WITH DIFFUSION AUTOENCODERS AND OFFLINE QUANTIZATION

ICASSP 2026oral

Neural audio compression models have recently achieved extreme compression rates, enabling efficient latent generative modeling. Conversely, latent generative models have been applied to compression, pushing the limits of continuous and discrete approaches. However, existing methods remain constrain…

Cited by 0SourcePDFScholar
2025

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

ICASSP 2025accepted

This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key attributes, such as speaker identity, pitch, content, loudness, signal-to-noise rat…

Cited by 0SourceScholar
2025

F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music Generation

ICASSP 2025accepted

While music remains a challenging domain for generative models like Transformers, recent progress has been made by exploiting suitable musically-informed priors. One technique to leverage information about musical structure in Transformers is inserting such knowledge into the positional encoding (PE…

Cited by 0SourceScholar
2025

Investigating the Sensitivity of Pre-trained Audio Embeddings to Common Effects

ICASSP 2025accepted

In recent years, foundation models have significantly advanced data-driven systems across various domains. Yet, their underlying properties, especially when functioning as feature extractors, remain under-explored. In this paper, we investigate the sensitivity to audio effects of audio embeddings ex…

Cited by 0SourceScholar
2025

Multiple Choice Learning for Efficient Speech Separation with Many Speakers

ICASSP 2025accepted

Training speech separation models in the supervised setting raises a permutation problem: finding the best assignation between the model predictions and the ground truth separated signals. This inherently ambiguous task is customarily solved using Permutation Invariant Training (PIT). In this articl…

Cited by 0SourceScholar
2024

A Fully Differentiable Model for Unsupervised Singing Voice Separation

ICASSP 2024accepted

A novel model was recently proposed by Schulze-Forster et al. in [1] for unsupervised music source separation. This model allows to tackle some of the major shortcomings of existing source separation frameworks. Specifically, it eliminates the need for isolated sources during training, performs effi…

Cited by 0SourceScholar
2024

Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing

NeurIPS 2024poster

We introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of plausible hypotheses. These hypotheses are trained using the Winner-takes-all (WTA) scheme, which promotes the diversit…

2024

GLA-GRAD: A Griffin-Lim Extended Waveform Generation Diffusion Model

ICASSP 2024accepted

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a diffusion process for the generation of high-fidelity audio. How…

Cited by 0SourceScholar
2024

SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis

ICASSP 2024accepted

Generative adversarial network (GAN) models can synthesize high-quality audio signals while ensuring fast sample generation. However, they are difficult to train and are prone to several issues including mode collapse and divergence. In this paper, we introduce SpecDiff-GAN, a neural vocoder based o…

Cited by 0SourceScholar
2024

Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport

ICASSP 2024accepted

In neural audio signal processing, pitch conditioning has been used to enhance the performance of synthesizers. However, jointly training pitch estimators and synthesizers is a challenge when using standard audio-to-audio reconstruction loss, leading to reliance on external pitch trackers. To addres…

Cited by 0SourceScholar
2024

Winner-takes-all learners are geometry-aware conditional density estimators

ICML 2024poster

Winner-takes-all training is a simple learning paradigm, which handles ambiguous tasks by predicting a set of plausible hypotheses. Recently, a connection was established between Winner-takes-all training and centroidal Voronoi tessellations, showing that, once trained, hypotheses should quantize op…

2023

Learning Interpretable Filters In Wav-UNet For Speech Enhancement

ICASSP 2023accepted

Due to their performances, deep neural networks have emerged as a major method in nearly all modern audio processing applications. Deep neural networks can be used to estimate some parameters or hyperparameters of a model, or in some cases the entire model in an end-to-end fashion. Although deep lea…

Cited by 0SourceScholar
2023

Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis

NeurIPS 2023poster

We introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may be sampled for each training input. Multiple Choice Learning is a simple framework to tackle multimodal density estimatio…

2022

Listen to Interpret: Post-hoc Interpretability for Audio Networks with NMF

NeurIPS 2022accept

This paper tackles post-hoc interpretability for audio processing networks. Our goal is to interpret decisions of a trained network in terms of high-level audio objects that are also listenable for the end-user. To this end, we propose a novel interpreter design that incorporates non-negative matrix…

2022

Phase Shifted Bedrosian Filterbank: An Interpretable Audio Front-End for Time-Domain Audio Source Separation

ICASSP 2022accepted

The use of a parameterized encoders or audio front-ends has shown promises in improving the interpretability of time domain single-channel source separation models such as Conv-TasNet. This type of filters also allows a potential reduction of the computational cost since larger encoder filters can b…

Cited by 0SourceScholar
2021

Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks

NeurIPS 2021poster

Neural network compression techniques have become increasingly popular as they can drastically reduce the storage and computation requirements for very large networks. Recent empirical studies have illustrated that even simple pruning strategies can be surprisingly effective, and several theoretical…

2021

Neuro-Steered Music Source Separation With EEG-Based Auditory Attention Decoding And Contrastive-NMF

ICASSP 2021accepted

We propose a novel informed music source separation paradigm, which can be referred to as neuro-steered music source separation. More precisely, the source separation process is guided by the user’s selective auditory attention decoded from his/her EEG response to the stimulus. This high-level prior…

Cited by 0SourceScholar
2020

Audio-Based Auto-Tagging With Contextual Tags for Music

ICASSP 2020accepted

Music listening context such as location or activity has been shown to greatly influence the users' musical tastes. In this work, we study the relationship between user context and audio content in order to enable context-aware music recommendation agnostic to user data. For that, we propose a semi-…

Cited by 0SourceScholar
2020

Audio-Based Detection of Explicit Content in Music

ICASSP 2020accepted

We present a novel automatic system for performing explicit content detection directly on the audio signal. Our modular approach uses an audio-to-character recognition model, a keyword spotting model associated with a dictionary of carefully chosen keywords, and a Random Forest classification model…

Cited by 0SourceScholar
2020

Joint Phoneme Alignment and Text-Informed Speech Separation on Highly Corrupted Speech

ICASSP 2020accepted

Speech separation quality can be improved by exploiting textual information. However, this usually requires text-to-speech alignment at phoneme level. Classical alignment methods are made for rather clean speech and do not work as well on corrupted speech. We propose to perform text-informed speech-…

Cited by 0SourceScholar
2020

Neutral to Lombard Speech Conversion with Deep Learning

ICASSP 2020accepted

In this paper, we propose several approaches for neutral to Lombard speech conversion. We study in particular the influence of different recurrent neural network architectures where their main hyper-parameters are carefully selected using a bandit-based approach. We also apply the Continuous Wavelet…

Cited by 2SourceScholar
2020

Speech Intelligibility Enhancement by Equalization for in-Car Applications

ICASSP 2020accepted

In this paper, we propose a speech intelligibility enhancement method for typical in-car applications in noisy environments. While traditional speech enhancement algorithms aim at increasing the Signal to Noise Ratio (SNR), the goal here is to increase intelligibility by applying dedicated voice tra…

Cited by 0SourceScholar
2019

First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise

NeurIPS 2019poster

Stochastic gradient descent (SGD) has been widely used in machine learning due to its computational efficiency and favorable generalization properties. Recently, it has been empirically demonstrated that the gradient noise in several deep learning settings admits a non-Gaussian, heavy-tailed behavio…

2018

Alpha-Stable Low-Rank Plus Residual Decomposition for Speech Enhancement

ICASSP 2018accepted

In this study, we propose a novel probabilistic model for separating clean speech signals from noisy mixtures by decomposing the mixture spectra into a structured speech part and a more flexible residual part. The main novelty in our model is that it uses a family of heavy-tailed distributions, so c…

Cited by 0SourceScholar
2017

Alpha-stable multichannel audio source separation

ICASSP 2017accepted

In this paper, we focus on modeling multichannel audio signals in the short-time Fourier transform domain for the purpose of source separation. We propose a probabilistic model based on a class of heavy-tailed distributions, in which the observed mixtures and the latent sources are jointly modeled b…

Cited by 0SourceScholar
2017

Drum extraction in single channel audio signals using multi-layer Non negative Matrix Factor Deconvolution

ICASSP 2017accepted

In this paper, we propose a supervised multilayer factorization method designed for harmonic/percussive source separation and drum extraction. Our method decomposes the audio signals in sparse orthogonal components which capture the harmonic content, while the drum is represented by an extension of…

Cited by 0SourceScholar
2017

Motion informed audio source separation

ICASSP 2017accepted

In this paper we tackle the problem of single channel audio source separation driven by descriptors of the sounding object's motion. As opposed to previous approaches, motion is included as a soft-coupling constraint within the nonnegative matrix factorization framework. The proposed method is appli…

Cited by 0SourceScholar
2017

Multichannel audio source separation: Variational inference of time-frequency sources from time-domain observations

ICASSP 2017accepted

A great number of methods for multichannel audio source separation are based on probabilistic approaches in which the sources are modeled as latent random variables in a Time-Frequency (TF) domain. For reverberant mixtures, it is common to approximate the time-domain convolutive mixing process as be…

Cited by 0SourceScholar
2017

Overlapping sound event detection with supervised Nonnegative Matrix Factorization

ICASSP 2017accepted

In this paper we propose a supervised Nonnegative Matrix Factorization (NMF) model for overlapping sound event detection in real life audio. We start by highlighting the usefulness of non-euclidean NMF to learn representations for detecting and classifying acoustic events in a multi-label setting. T…

Cited by 0SourceScholar
2017

Parallelized Stochastic Gradient Markov Chain Monte Carlo algorithms for non-negative matrix factorization

ICASSP 2017accepted

Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have become popular in modern data analysis problems due to their computational efficiency. Even though they have proved useful for many statistical models, the application of SG-MCMC to non-negative matrix factorization (NMF) models has…

Cited by 0SourceScholar
2017

Supervised group nonnegative matrix factorisation with similarity constraints and applications to speaker identification

ICASSP 2017accepted

This paper presents supervised feature learning approaches for speaker identification that rely on nonnegative matrix factorisation. Recent studies have shown that group nonnegative matrix factorisation and task-driven supervised dictionary learning can help performing effective feature learning for…

Cited by 0SourceScholar
2016

Acoustic scene classification with matrix factorization for unsupervised feature learning

ICASSP 2016accepted

In this paper we study the use of unsupervised feature learning for acoustic scene classification (ASC). The acoustic environment recordings are represented by time-frequency images from which we learn features in an unsupervised manner. After a set of preprocessing and pooling steps, the images are…

Cited by 0SourceScholar
2016

Feature adapted convolutional neural networks for downbeat tracking

ICASSP 2016accepted

We define a novel system for the automatic estimation of downbeat positions from audio music signals. New rhythm and melodic features are introduced and feature adapted convolutional neural networks are used to take advantage of their specificity. Indeed, invariance to melody transposition, chroma d…

Cited by 23SourceScholar
2016

Formant shifting for speech intelligibility improvement in car noise environment

ICASSP 2016accepted

In this paper, we propose a novel approach aiming at improving the intelligibility of speech in the context of in-car applications. Speech produced in noisy environments is subject to the Lombard effect which gathers a number of voice transformation effects compared to the speech produced in calm en…

Cited by 0SourceScholar
2016

Group nonnegative matrix factorisation with speaker and session variability compensation for speaker identification

ICASSP 2016accepted

This paper presents a feature learning approach for speaker identification that is based on nonnegative matrix factorisation. Recent studies have shown that with such models, the dictionary atoms can represent well the speaker identity. The approaches proposed so far focused only on speaker variabil…

Cited by 0SourceScholar
2016

Stochastic Gradient Richardson-Romberg Markov Chain Monte Carlo

NeurIPS 2016poster

Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) algorithms have become increasingly popular for Bayesian inference in large-scale applications. Even though these methods have proved useful in several scenarios, their performance is often limited by their bias. In this study, we propose a nove…

Cited by 42SourcePDFScholar
2016

Stochastic thermodynamic integration: Efficient Bayesian model selection via stochastic gradient MCMC

ICASSP 2016accepted

Model selection is a central topic in Bayesian machine learning, which requires the estimation of the marginal likelihood of the data under the models to be compared. During the last decade, conventional model selection methods have lost their charm as they have high computational requirements. In t…

Cited by 0SourceScholar
2015

Downbeat tracking with multiple features and deep neural networks

ICASSP 2015accepted

In this paper, we introduce a novel method for the automatic estimation of downbeat positions from music signals. Our system relies on the computation of musically inspired features capturing important aspects of music such as timbre, harmony, rhythmic patterns, or local similarities in both timbre…

Cited by 39SourceScholar
2015

Multipitch estimation using a PLCA-based model: Impact of partial user annotation

ICASSP 2015accepted

In this paper one investigates the merit of partial user annotation for music transcription using a PLCA-based model. The original algorithm, called Blind Harmonic Adaptive Decomposition (BHAD), provides an estimation of the polyphonic pitch content of the input signal in an entirely unsupervised ma…

Cited by 0SourceScholar