← Search

Mathieu Fontaine

13 accepted papers

2026

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

ICML 2026poster

We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference time. Traditional language modeling is an intrinsically ill-posed problem: given a context, multiple ``futures'' may be…

Cited by 0SourceScholar
2026

SIRUP: A DIFFUSION-BASED VIRTUAL UPMIXER OF STEERING VECTORS FOR HIGHLY-DIRECTIVE SPATIALIZATION WITH FIRST-ORDER AMBISONICS

ICASSP 2026poster

This paper presents virtual upmixing of steering vectors captured by a fewer-channel spherical microphone array. This challenge has conventionally been addressed by recovering the directions and signals of sound sources from first-order ambisonics (FOA) data, and then rendering the higher-order ambi…

Cited by 0SourcePDFScholar
2025

Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement

ICASSP 2025accepted

Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of…

Cited by 0SourceScholar
2025

O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization

ICASSP 2025accepted

We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we develop a novel centroid refinement decoder whose usefulness is assessed through a rigorous ablation study. Our system p…

Cited by 0SourceScholar
2024

GLA-GRAD: A Griffin-Lim Extended Waveform Generation Diffusion Model

ICASSP 2024accepted

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a diffusion process for the generation of high-fidelity audio. How…

Cited by 0SourceScholar
2024

SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis

ICASSP 2024accepted

Generative adversarial network (GAN) models can synthesize high-quality audio signals while ensuring fast sample generation. However, they are difficult to train and are prone to several issues including mode collapse and divergence. In this paper, we introduce SpecDiff-GAN, a neural vocoder based o…

Cited by 0SourceScholar
2024

Winner-takes-all learners are geometry-aware conditional density estimators

ICML 2024poster

Winner-takes-all training is a simple learning paradigm, which handles ambiguous tasks by predicting a set of plausible hypotheses. Recently, a connection was established between Winner-takes-all training and centroidal Voronoi tessellations, showing that, once trained, hypotheses should quantize op…

2023

Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis

NeurIPS 2023poster

We introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may be sampled for each training input. Multiple Choice Learning is a simple framework to tackle multimodal density estimatio…

2022

Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational Environments

IROS 2022poster

This paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g., cocktail party). One may use a state-of-the-art blind source sepa…

Cited by 6SourceScholar
2022

Flow-Based Fast Multichannel Nonnegative Matrix Factorization for Blind Source Separation

ICASSP 2022accepted

This paper describes a blind source separation method for multichannel audio signals, called NF-FastMNMF, based on the integration of the normalizing flow (NF) into the multichannel nonnegative matrix factorization with jointly-diagonalizable spatial covariance matrices, a.k.a. FastMNMF. Whereas the…

Cited by 0SourceScholar
2021

Autoregressive Fast Multichannel Nonnegative Matrix Factorization For Joint Blind Source Separation And Dereverberation

ICASSP 2021accepted

This paper describes a joint blind source separation and dereverberation method that works adaptively and efficiently in a reverberant noisy environment. The modern approach to blind source separation (BSS) is to formulate a probabilistic model of multichannel mixture signals that consists of a sour…

Cited by 0SourceScholar