← Search

Jean-François Bonastre

11 accepted papers

2026

SEGMENTWISE PRUNING IN AUDIO-LANGUAGE MODELS

ICASSP 2026poster

Recent audio-language models have shown impressive performance across a wide range of audio tasks and are increasingly capable of handling long audio inputs. However, the computing costs in these models heavily depend on sequence length, which can become very large given the nature of audio data. In…

Cited by 0SourcePDFScholar
2024

Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset

ICASSP 2024accepted

The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech coll…

Cited by 0SourceScholar
2023

Federated Learning for ASR Based on wav2vec 2.0

ICASSP 2023accepted

This paper presents a study on the use of federated learning to train an ASR model based on a wav2vec 2.0 model pre-trained by self supervision. Carried out on the well-known TED-LIUM 3 dataset, our experiments show that such a model can obtain, with no use of a language model, a word error rate of…

Cited by 0SourceScholar
2023

Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline

ICASSP 2023accepted

The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representatio…

Cited by 0SourceScholar
2022

A Bridge between Features and Evidence for Binary Attribute-Driven Perfect Privacy

ICASSP 2022accepted

Attribute-driven privacy aims to conceal a single user’s attribute, contrary to anonymisation that tries to hide the full identity of the user in some data. When the attribute to protect from malicious inferences is binary, perfect privacy requires the log-likelihood-ratio to be zero resulting in no…

Cited by 0SourceScholar
2022

Privacy Attacks for Automatic Speech Recognition Acoustic Models in A Federated Learning Framework

ICASSP 2022accepted

This paper investigates methods to effectively retrieve speaker information from the personalized speaker adapted neural network acoustic models (AMs) in automatic speech recognition (ASR). This problem is especially important in the context of federated learning of ASR acoustic models where a globa…

Cited by 0SourceScholar
2022

Retrieving Speaker Information from Personalized Acoustic Models for Speech Recognition

ICASSP 2022accepted

The widespread of powerful personal devices capable of collecting voice of their users has opened the opportunity to build speaker adapted speech recognition system (ASR) or to participate to collaborative learning of ASR. In both cases, personalized acoustic models (AM), i.e. fine-tuned AM with spe…

Cited by 0SourceScholar
2019

Similarity Metric Based on Siamese Neural Networks for Voice Casting

ICASSP 2019accepted

Dubbing contributes to a larger international distribution of multimedia documents. It aims to replace the original voice in a source language by a new one in a target language. For now, the target voice selection procedure, called voice casting, is manually performed by human experts. This selectio…

Cited by 0SourceScholar
2017

Phonological content impact on wrongful convictions in Forensic Voice Comparison context

ICASSP 2017accepted

Forensic Voice Comparison (FVC) is increasingly using the likelihood ratio (LR) in order to indicate whether the evidence supports the prosecution (same-speaker) or defender (different-speakers) hypotheses. Nevertheless, the LR accepts some practical limitations due both to its estimation process it…

Cited by 0SourceScholar
2016

Inter-speaker variability in forensic voice comparison: A preliminary evaluation

ICASSP 2016accepted

In forensic voice comparison, it is strongly recommended to follow Bayesian paradigm. In this paradigm, the strength of the forensic evidence is summarized by a likelihood ratio (LR). The LR magnitude quantifies the strength of the evidence: far from unity for a meaningful LR (a LR which supports st…

Cited by 0SourceScholar
2015

Additive noise compensation in the i-vector space for speaker recognition

ICASSP 2015accepted

State-of-the-art speaker recognition systems performance degrades considerably in noisy environments even though they achieve very good results in clean conditions. In order to deal with this strong limitation, we aim in this work to remove the noisy part of an i-vector directly in the i-vector spac…

Cited by 0SourceScholar