← Search

Slim Essid

28 accepted papers

2026

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

ICML 2026poster

We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference time. Traditional language modeling is an intrinsically ill-posed problem: given a context, multiple ``futures'' may be…

Cited by 0SourceScholar
2025

Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement

ICASSP 2025accepted

Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of…

Cited by 0SourceScholar
2025

Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning

ICASSP 2025accepted

Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed to extract higher-level information that could be more suited for down-stream cla…

Cited by 0SourceScholar
2025

Multiple Choice Learning for Efficient Speech Separation with Many Speakers

ICASSP 2025accepted

Training speech separation models in the supervised setting raises a permutation problem: finding the best assignation between the model predictions and the ground truth separated signals. This inherently ambiguous task is customarily solved using Permutation Invariant Training (PIT). In this articl…

Cited by 0SourceScholar
2025

O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization

ICASSP 2025accepted

We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we develop a novel centroid refinement decoder whose usefulness is assessed through a rigorous ablation study. Our system p…

Cited by 0SourceScholar
2025

Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping

ICASSP 2025accepted

People often listen to music in noisy environments, seeking to isolate themselves from ambient sounds. Indeed, a music signal can mask some of the noise’s frequency components due to the effect of simultaneous masking. In this article, we propose a neural network based on a psychoacoustic masking mo…

Cited by 0SourceScholar
2024

Adapting Pitch-Based Self Supervised Learning Models for Tempo Estimation

ICASSP 2024accepted

Tempo estimation is the task of estimating the periodicity of the dominant rhythm pulse of a music audio signal. It has therefore a close relationship with dominant pitch estimation. Recently, both tasks have been addressed in a Self-Supervised Learning (SSL) fashion so as to leverage unlabelled dat…

Cited by 0SourceScholar
2024

An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching

NeurIPS 2024poster

Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In this work, we show that this ability can be re-purposed for audio captioning, where the joint image-language decoder can b…

Cited by 1SourcePDFScholar
2024

Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing

NeurIPS 2024poster

We introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of plausible hypotheses. These hypotheses are trained using the Winner-takes-all (WTA) scheme, which promotes the diversit…

2024

Collaborating Foundation Models for Domain Generalized Semantic Segmentation

CVPR 2024poster

Domain Generalized Semantic Segmentation (DGSS) deals with training a model on a labeled source domain with the aim of generalizing to unseen domains during inference. Existing DGSS methods typically effectuate robust features by means of Domain Randomization (DR). Such an approach is often limited…

2024

On The Choice of the Optimal Temporal Support for Audio Classification with Pre-Trained Embeddings

ICASSP 2024accepted

Current state-of-the-art audio analysis systems rely on pre-trained embedding models, often used off-the-shelf as (frozen) feature extractors. Choosing the best one for a set of tasks is the subject of many recent publications. However, one aspect often overlooked in these works is the influence of…

Cited by 0SourceScholar
2024

Winner-takes-all learners are geometry-aware conditional density estimators

ICML 2024poster

Winner-takes-all training is a simple learning paradigm, which handles ambiguous tasks by predicting a set of plausible hypotheses. Recently, a connection was established between Winner-takes-all training and centroidal Voronoi tessellations, showing that, once trained, hypotheses should quantize op…

2023

Cosmopolite Sound Monitoring (CoSMo): A Study of Urban Sound Event Detection Systems Generalizing to Multiple Cities

ICASSP 2023accepted

Measuring noise in cities and automatically identifying the corresponding sound sources are a crucial challenge for policymakers. Indeed, such information helps addressing noise pollution and improving the well-being of urban dwellers. In recent years, researchers have provided annotated datasets re…

Cited by 0SourceScholar
2023

Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis

NeurIPS 2023poster

We introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may be sampled for each training input. Multiple Choice Learning is a simple framework to tackle multimodal density estimatio…

2021

Distributed Speech Separation in Spatially Unconstrained Microphone Arrays

ICASSP 2021accepted

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different sources using sophisticated deep neural networks which are very…

Cited by 0SourceScholar
2021

Neuro-Steered Music Source Separation With EEG-Based Auditory Attention Decoding And Contrastive-NMF

ICASSP 2021accepted

We propose a novel informed music source separation paradigm, which can be referred to as neuro-steered music source separation. More precisely, the source separation process is guided by the user’s selective auditory attention decoded from his/her EEG response to the stimulus. This high-level prior…

Cited by 0SourceScholar
2020

DNN-based Distributed Multichannel Mask Estimation for Speech Enhancement in Microphone Arrays

ICASSP 2020accepted

Multichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions in the real world. Distributed sensor arrays that consider several devices with a few microphones is a viable solution which allows for exploiting the multiple devices e…

Cited by 0SourceScholar
2019

A Music Structure Informed Downbeat Tracking System Using Skip-chain Conditional Random Fields and Deep Learning

ICASSP 2019accepted

In recent years the task of downbeat tracking has received increasing attention and the state of the art has been improved with the introduction of deep learning methods. Among proposed solutions, existing systems exploit short-term musical rules as part of their language modelling. In this work we…

Cited by 0SourceScholar
2018

An Ensemble Learning Approach to Detect Epileptic Seizures from Long Intracranial EEG Recordings

ICASSP 2018accepted

This paper proposes a patient-specific supervised classification algorithm to detect seizures in long offline intracranial electroencephalographic (iEEG) recordings. The main idea of the proposed algorithm is to combine a set of probabilistic classifiers, trained on a dataset of 1 s epochs, into a w…

Cited by 0SourceScholar
2018

Attitude Classification in Adjacency Pairs of a Human-Agent Interaction with Hidden Conditional Random Fields

ICASSP 2018accepted

In this paper, the main goal is to classify, in a human-agent interaction, the attitude of the user using hidden conditional random fields. This model allows us to capture the dynamics of the interaction in the pairs of speech turns (adjacency pairs) analyzed by our system. High level linguistic fea…

Cited by 0SourceScholar
2018

Structured Output Learning with Abstention: Application to Accurate Opinion Prediction

ICML 2018oral

Motivated by Supervised Opinion Analysis, we propose a novel framework devoted to Structured Output Learning with Abstention (SOLA). The structure prediction model is able to abstain from predicting some labels in the structured output at a cost chosen by the user in a flexible way. For that purpose…

Cited by 5SourcePDFScholar
2017

Motion informed audio source separation

ICASSP 2017accepted

In this paper we tackle the problem of single channel audio source separation driven by descriptors of the sounding object's motion. As opposed to previous approaches, motion is included as a soft-coupling constraint within the nonnegative matrix factorization framework. The proposed method is appli…

Cited by 0SourceScholar
2017

Overlapping sound event detection with supervised Nonnegative Matrix Factorization

ICASSP 2017accepted

In this paper we propose a supervised Nonnegative Matrix Factorization (NMF) model for overlapping sound event detection in real life audio. We start by highlighting the usefulness of non-euclidean NMF to learn representations for detecting and classifying acoustic events in a multi-label setting. T…

Cited by 0SourceScholar
2017

Supervised group nonnegative matrix factorisation with similarity constraints and applications to speaker identification

ICASSP 2017accepted

This paper presents supervised feature learning approaches for speaker identification that rely on nonnegative matrix factorisation. Recent studies have shown that group nonnegative matrix factorisation and task-driven supervised dictionary learning can help performing effective feature learning for…

Cited by 0SourceScholar
2016

Acoustic scene classification with matrix factorization for unsupervised feature learning

ICASSP 2016accepted

In this paper we study the use of unsupervised feature learning for acoustic scene classification (ASC). The acoustic environment recordings are represented by time-frequency images from which we learn features in an unsupervised manner. After a set of preprocessing and pooling steps, the images are…

Cited by 0SourceScholar
2016

Group nonnegative matrix factorisation with speaker and session variability compensation for speaker identification

ICASSP 2016accepted

This paper presents a feature learning approach for speaker identification that is based on nonnegative matrix factorisation. Recent studies have shown that with such models, the dictionary atoms can represent well the speaker identity. The approaches proposed so far focused only on speaker variabil…

Cited by 0SourceScholar