← Search

Romain Serizel

32 accepted papers

2026

Domain-Invariant Representation Learning of Bird Sounds

ICASSP 2026poster

Passive acoustic monitoring (PAM) is crucial for bioacoustic research, enabling non-invasive species tracking and biodiversity monitoring. Citizen science platforms provide large annotated datasets from focal recordings, where the target species is intentionally recorded. However, PAM requires monit…

Cited by 0SourcePDFScholar
2025

A decade of DCASE: Achievements, practices, evaluations and future challenges

ICASSP 2025accepted

This paper introduces briefly the history and growth of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, workshop, research area and research community. Created in 2013 as a data evaluation challenge, DCASE has become a major research topic in the Audio and Acoustic…

Cited by 0SourceScholar
2025

Diffusion-based Unsupervised Audio-visual Speech Enhancement

ICASSP 2025accepted

This paper proposes a new unsupervised audiovisual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion model is pre-trained on clean speech conditioned on correspondin…

Cited by 0SourceScholar
2025

Latent Watermarking of Audio Generative Models

ICASSP 2025accepted

The advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. To address this, watermarking techniques have been recently developed, enabling the detection of content generated by a deployed model. For such techniques to…

Cited by 0SourceScholar
2024

A Weighted-Variance Variational Autoencoder Model for Speech Enhancement

ICASSP 2024accepted

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed for the generative model, where the speech information is encoded in the varianc…

Cited by 0SourceScholar
2024

Diffusion-Based Speech Enhancement with a Weighted Generative-Supervised Learning Loss

ICASSP 2024accepted

Diffusion-based generative models have recently gained attention in speech enhancement (SE), providing an alternative to conventional supervised methods. These models transform clean speech training samples into Gaussian noise, usually centered on noisy speech, and subsequently learn a parameterized…

Cited by 0SourceScholar
2024

Performance and Energy Balance: A Comprehensive Study of State-of-the-Art Sound Event Detection Systems

ICASSP 2024accepted

In recent years, deep learning systems have shown a concerning trend toward increased complexity and higher energy consumption. As researchers in this domain and organizers of one of the Detection and Classification of Acoustic Scenes and Events challenges task, we recognize the importance of addres…

Cited by 0SourceScholar
2024

Posterior Sampling Algorithms for Unsupervised Speech Enhancement with Recurrent Variational Autoencoder

ICASSP 2024accepted

In this paper, we address the unsupervised speech enhancement problem based on recurrent variational autoencoder (RVAE). This approach offers promising generalization performance over the supervised counterpart. Nevertheless, the involved iterative variational expectation-maximization (VEM) process…

Cited by 0SourceScholar
2024

RoboVox: A Single/Multi-channel Far-field Speaker Recognition Benchmark for a Mobile Robot

COLING 2024main

In this paper, we introduce a new far-field speaker recognition benchmark called RoboVox. RoboVox is a French corpus recorded by a mobile robot. The files are recorded from different distances under severe acoustical conditions with the presence of several types of noise and reverberation. In additi…

Cited by 0SourcePDFScholar
2024

Unsupervised Speech Enhancement with Diffusion-Based Generative Models

ICASSP 2024accepted

Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when generalising to unseen conditions. To address this issue, we introduce an alterna…

Cited by 0SourceScholar
2023

Audio-Visual Speech Enhancement with a Deep Kalman Filter Generative Model

ICASSP 2023accepted

Deep latent variable generative models based on variational autoencoder (VAE) have shown promising performance for audio-visual speech enhancement (AVSE). The underlying idea is to learn a VAE-based audio-visual prior distribution for clean speech data, and then combine it with a statistical noise m…

Cited by 0SourceScholar
2023

From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion

NeurIPS 2023poster

Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthesize audio waveforms conditioned on highly compressed representations. Although…

Cited by 25SourcePDFScholar
2023

Lightweight Annotation and Class Weight Training for Automatic Estimation of Alarm Audibility in Noise

ICASSP 2023accepted

In an effort to improve occupational health and safety, we recently proposed an approach to assess the audibility of acoustic danger signals. It is based on the use of a binary classifier trained on perceptual data to predict the audibility of acoustic alarms in audio clips. In the present article,…

Cited by 0SourceScholar
2023

Spice+: Evaluation of Automatic Audio Captioning Systems with Pre-Trained Language Models

ICASSP 2023accepted

Audio captioning aims at describing acoustic scenes with natural language. Systems are currently evaluated by image captioning metrics CIDEr and SPICE. However, recent studies have highlighted a poor correlation of these metrics with human assessments. In this paper, we propose SPICE+, a modificatio…

Cited by 0SourceScholar
2022

A Benchmark of State-of-the-Art Sound Event Detection Systems Evaluated on Synthetic Soundscapes

ICASSP 2022accepted

This paper proposes a benchmark of submissions to Detection and Classification Acoustic Scene and Events 2021 Challenge (DCASE) Task 4 representing a sampling of the state-of-the-art in Sound Event Detection task. The submissions are evaluated according to the two polyphonic sound detection score sc…

Cited by 0SourceScholar
2021

Distributed Speech Separation in Spatially Unconstrained Microphone Arrays

ICASSP 2021accepted

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different sources using sophisticated deep neural networks which are very…

Cited by 0SourceScholar
2021

Improving Sound Event Detection Metrics: Insights from DCASE 2020

ICASSP 2021accepted

The ranking of sound event detection (SED) systems may be biased by assumptions inherent to evaluation criteria and to the choice of an operating point. This paper compares conventional event-based and segment-based criteria against the Polyphonic Sound Detection Score (PSDS)'s intersection-based cr…

Cited by 0SourceScholar
2021

Sound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes

ICASSP 2021accepted

We propose a benchmark of state-of-the-art sound event detection systems (SED). We design synthetic evaluation sets to focus on specific sound event detection challenges. We analyze the performance of the submissions to DCASE 2020 Task 4 as a function of time-related modifications (time position of…

Cited by 0SourceScholar
2021

What's all the Fuss about Free Universal Sound Separation Data?

ICASSP 2021accepted

We introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio data drawn from 357 classes, which are used to create mixtur…

Cited by 0SourceScholar
2020

DNN-based Distributed Multichannel Mask Estimation for Speech Enhancement in Microphone Arrays

ICASSP 2020accepted

Multichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions in the real world. Distributed sensor arrays that consider several devices with a few microphones is a viable solution which allows for exploiting the multiple devices e…

Cited by 0SourceScholar
2020

Sound Event Detection in Synthetic Domestic Environments

ICASSP 2020accepted

We present a comparative analysis of the performance of state-of-the-art sound event detection systems. In particular, we study the robustness of the systems to noise and signal degradation, which is known to impact model generalization. Our analysis is based on the results of task 4 of the DCASE 20…

Cited by 0SourceScholar
2019

Semi-supervised Triplet Loss Based Learning of Ambient Audio Embeddings

ICASSP 2019accepted

Deep neural networks are particularly useful to learn relevant representations from data. Recent studies have demonstrated the potential of unsupervised representation learning for ambient sound analysis using various flavors of the triplet loss. They have compared this approach to supervised learni…

Cited by 0SourceScholar
2018

Multichannel Speech Separation with Recurrent Neural Networks from High-Order Ambisonics Recordings

ICASSP 2018accepted

We present a source separation system for high-order ambisonics (HOA) contents. We derive a multichannel spatial filter from a mask estimated by a long short-term memory (LSTM) recurrent neural network. We combine one channel of the mixture with the outputs of basic HOA beamformers as inputs to the…

Cited by 0SourceScholar
2018

Multiple-Input Neural Network-Based Residual Echo Suppression

ICASSP 2018accepted

A residual echo suppressor (RES) aims to suppress the residual echo in the output of an acoustic echo canceler (AEC). Spectral-based RES approaches typically estimate the magnitude spectra of the near-end speech and the residual echo from a single input, that is either the far-end speech or the echo…

Cited by 0SourceScholar
2017

Supervised group nonnegative matrix factorisation with similarity constraints and applications to speaker identification

ICASSP 2017accepted

This paper presents supervised feature learning approaches for speaker identification that rely on nonnegative matrix factorisation. Recent studies have shown that group nonnegative matrix factorisation and task-driven supervised dictionary learning can help performing effective feature learning for…

Cited by 0SourceScholar
2016

Acoustic scene classification with matrix factorization for unsupervised feature learning

ICASSP 2016accepted

In this paper we study the use of unsupervised feature learning for acoustic scene classification (ASC). The acoustic environment recordings are represented by time-frequency images from which we learn features in an unsupervised manner. After a set of preprocessing and pooling steps, the images are…

Cited by 0SourceScholar
2016

Group nonnegative matrix factorisation with speaker and session variability compensation for speaker identification

ICASSP 2016accepted

This paper presents a feature learning approach for speaker identification that is based on nonnegative matrix factorisation. Recent studies have shown that with such models, the dictionary atoms can represent well the speaker identity. The approaches proposed so far focused only on speaker variabil…

Cited by 0SourceScholar