← Search

Emmanuel Vincent

24 accepted papers

2026

BEST-RQ-BASED SELF-SUPERVISED LEARNING FOR WHISPER DOMAIN ADAPTATION

ICASSP 2026poster

Automatic Speech Recognition (ASR) systems, despite large multilingual training, struggle in low-resource scenarios where labeled data is scarce. We propose BEARD (BEST-RQ Encoder Adaptation with Re-training and Distillation), a novel framework designed to adapt Whisper's encoder with unlabeled data…

Cited by 0SourcePDFScholar
2025

Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization

ICASSP 2025accepted

In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verification based only on phoneme durations. Experimental results demonstrate that pho…

Cited by 0SourceScholar
2024

MMAR: Multilingual and Multimodal Anaphora Resolution in Instructional Videos

EMNLP 2024finding

Multilingual anaphora resolution identifies referring expressions and implicit arguments in texts and links to antecedents that cover several languages. In the most challenging setting, cross-lingual anaphora resolution, training data, and test data are in different languages. As knowledge needs to…

2023

Find-2-Find: Multitask Learning for Anaphora Resolution and Object Localization

EMNLP 2023long main

In multimodal understanding tasks, visual and linguistic ambiguities can arise. Visual ambiguity can occur when visual objects require a model to ground a referring expression in a video without strong supervision, while linguistic ambiguity can occur from changes in entities in action flows. As an…

Cited by 0SourceScholar
2022

On The Impact of Normalization Strategies in Unsupervised Adversarial Domain Adaptation for Acoustic Scene Classification

ICASSP 2022accepted

Acoustic scene classification systems face performance degradation when trained and tested on data recorded by different devices. Unsupervised domain adaptation methods have been studied to reduce the impact of this mismatch. While they do not assume the availability of labels at test time, they oft…

Cited by 0SourceScholar
2020

Evaluating Voice Conversion-Based Privacy Protection against Informed Attackers

ICASSP 2020accepted

Speech data conveys sensitive speaker attributes like identity or accent. With a small amount of found data, such attributes can be inferred and exploited for malicious purposes: voice cloning, spoofing, etc. Anonymization aims to make the data unlinkable, i.e., ensure that no utterance can be linke…

Cited by 0SourceScholar
2020

Filterbank Design for End-to-end Speech Separation

ICASSP 2020accepted

Single-channel speech separation has recently made great progress thanks to learned filterbanks as used in ConvTasNet. In parallel, parameterized filterbanks have been proposed for speaker recognition where only center frequencies and bandwidths are learned. In this work, we extend real-valued learn…

Cited by 0SourceScholar
2020

SLOGD: Speaker Location Guided Deflation Approach to Speech Separation

ICASSP 2020accepted

Speech separation is the process of separating multiple speakers from an audio recording. In this work we propose to separate the sources using a Speaker LOcalization Guided Deflation (SLOGD) approach wherein we estimate the sources iteratively. In each iteration we first estimate the location of th…

Cited by 3SourceScholar
2019

An Improved Uncertainty Propagation Method for Robust I-vector Based Speaker Recognition

ICASSP 2019accepted

The performance of automatic speaker recognition systems degrades when facing distorted speech data containing additive noise and/or reverberation. Statistical uncertainty propagation has been introduced as a promising paradigm to address this challenge. So far, different uncertainty propagation met…

Cited by 0SourceScholar
2019

Semi-supervised Triplet Loss Based Learning of Ambient Audio Embeddings

ICASSP 2019accepted

Deep neural networks are particularly useful to learn relevant representations from data. Recent studies have demonstrated the potential of unsupervised representation learning for ambient sound analysis using various flavors of the triplet loss. They have compared this approach to supervised learni…

Cited by 0SourceScholar
2018

Multichannel Speech Separation with Recurrent Neural Networks from High-Order Ambisonics Recordings

ICASSP 2018accepted

We present a source separation system for high-order ambisonics (HOA) contents. We derive a multichannel spatial filter from a mask estimated by a long short-term memory (LSTM) recurrent neural network. We combine one channel of the mixture with the outputs of basic HOA beamformers as inputs to the…

Cited by 43SourceScholar
2018

Multiple-Input Neural Network-Based Residual Echo Suppression

ICASSP 2018accepted

A residual echo suppressor (RES) aims to suppress the residual echo in the output of an acoustic echo canceler (AEC). Spectral-based RES approaches typically estimate the magnitude spectra of the near-end speech and the residual echo from a single input, that is either the far-end speech or the echo…

Cited by 0SourceScholar
2018

Semi-Supervised Learning with Deep Neural Networks for Relative Transfer Function Inverse Regression

ICASSP 2018accepted

Prior knowledge of the relative transfer function (RTF) is useful in many applications but remains little studied. In this paper, we propose a semi-supervised learning algorithm based on deep neural networks (DNNs) for RTF inverse regression, that is to generate the full-band RTF vector directly fro…

Cited by 0SourceScholar
2017

Discriminative importance weighting of augmented training data for acoustic model training

ICASSP 2017accepted

DNN based acoustic models require a large amount of training data. Parametric data augmentation techniques such as adding noise, reverberation, or changing the speech rate, are often employed to boost the dataset size and the ASR performance. The choice of augmentation techniques and the associated…

Cited by 0SourceScholar
2017

Recursive Bayesian estimation of the acoustic noise emitted by wind farms

ICASSP 2017accepted

Wind turbine noise is often annoying for humans living in close proximity to a wind farm. Reliably estimating the intensity of wind turbine noise is a necessary step towards quantifying and reducing annoyance, but it is challenging because of the overlap with background noise sources. Current approa…

Cited by 0SourceScholar
2016

Localizing an intermittent and moving sound source using a mobile robot

IROS 2016poster

This paper addresses the problem of localizing and tracking one intermittent, moving sound source using a microphone array on a mobile robot. Robot motion provides a solution for estimating the distance to the source and avoiding front-back ambiguity. We propose a mixture Kalman filter (MKF) framewo…

Cited by 23SourceScholar
2015

Micbots: Collecting large realistic datasets for speech and audio research using mobile robots

ICASSP 2015accepted

Speech and audio signal processing research is a tale of data collection efforts and evaluation campaigns. Large benchmark datasets for automatic speech recognition (ASR) have been instrumental in the advancement of speech recognition technologies. However, when it comes to robust ASR, source separa…

Cited by 0SourceScholar
2015

Music separation guided by cover tracks: Designing the joint NMF model

ICASSP 2015accepted

In audio source separation, reference guided approaches are a class of methods that use reference signals to guide the separation. In prior work, we proposed a general framework to model the deformation between the sources and the references. In this paper, we investigate a specific scenario within…

Cited by 0SourceScholar