← Search

Yannick Estève

14 accepted papers

2026

A STUDY OF DATA SELECTION STRATEGIES FOR PRE-TRAINING SELF-SUPERVISED SPEECH MODELS

ICASSP 2026oral

Self-supervised learning (SSL) has transformed speech processing, yet its reliance on massive pre-training datasets remains a bottleneck. While robustness is often attributed to scale and diversity, the role of the data distribution is less understood. We systematically examine how curated subsets o…

Cited by 0SourcePDFScholar
2026

Simultaneous Speech-to-Speech Translation Without Aligned Data

ICML 2026oral

Simultaneous speech translation is the task of translating source speech into a target language in real-time. Given that the dependencies between source and target words are non-monotonic (e.g. the word order can change between German and English), this means learning to jointly align and translate.…

Cited by 0SourcecodeScholar
2024

Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants

COLING 2024main

Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets with controlled demographic tags. This paper introduces the Sonos Voice Control B…

Cited by 1SourcePDFScholar
2024

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

COLING 2024main

In recent years, there has been a significant increase in interest in developing Spoken Language Understanding (SLU) systems. SLU involves extracting a list of semantic information from the speech signal. A major issue for SLU systems is the lack of sufficient amount of bi-modal (audio and textual s…

2023

Federated Learning for ASR Based on wav2vec 2.0

ICASSP 2023accepted

This paper presents a study on the use of federated learning to train an ASR model based on a wav2vec 2.0 model pre-trained by self supervision. Carried out on the well-known TED-LIUM 3 dataset, our experiments show that such a model can obtain, with no use of a language model, a word error rate of…

Cited by 0SourceScholar
2022

Privacy Attacks for Automatic Speech Recognition Acoustic Models in A Federated Learning Framework

ICASSP 2022accepted

This paper investigates methods to effectively retrieve speaker information from the personalized speaker adapted neural network acoustic models (AMs) in automatic speech recognition (ASR). This problem is especially important in the context of federated learning of ASR acoustic models where a globa…

Cited by 0SourceScholar
2022

Retrieving Speaker Information from Personalized Acoustic Models for Speech Recognition

ICASSP 2022accepted

The widespread of powerful personal devices capable of collecting voice of their users has opened the opportunity to build speaker adapted speech recognition system (ASR) or to participate to collaborative learning of ASR. In both cases, personalized acoustic models (AM), i.e. fine-tuned AM with spe…

Cited by 0SourceScholar
2021

An Empirical Study of End-To-End Simultaneous Speech Translation Decoding Strategies

ICASSP 2021accepted

This paper proposes a decoding strategy for end-to-end simultaneous speech translation. We leverage end-to-end models trained in offline mode and conduct an empirical study for two language pairs (English-to-German and English-to-Portuguese). We also investigate different output token granularities…

Cited by 0SourceScholar
2021

End2End Acoustic to Semantic Transduction

ICASSP 2021accepted

In this paper, we propose a novel end-to-end sequence-to-sequence spoken language understanding model using an attention mechanism. It reliably selects contextual acoustic features in order to hypothesize semantic contents. An initial architecture capable of extracting all pronounced words and conce…

Cited by 0SourceScholar
2021

Task Agnostic and Task Specific Self-Supervised Learning from Speech with LeBenchmark

NeurIPS 2021poster

Self-Supervised Learning (SSL) has yielded remarkable improvements in many different domains including computer vision, natural language processing and speech processing by leveraging large amounts of unlabeled data. In the specific context of speech, however, and despite promising results, there ex…

Cited by 41SourceScholar
2020

Dialogue History Integration into End-to-End Signal-to-Concept Spoken Language Understanding Systems

ICASSP 2020accepted

This work investigates the embeddings for representing dialog history in spoken language understanding (SLU) systems. We focus on the scenario when the semantic information is extracted directly from the speech signal by means of a single end-to-end neural network model. We proposed to integrate dia…

Cited by 15SourceScholar
2020

Error Analysis Applied to End-to-End Spoken Language Understanding

ICASSP 2020accepted

This paper presents a qualitative study of errors produced by an end-to-end spoken language understanding (SLU) system (speech signal to concepts) that reaches state of the art performance. Different studies are proposed to better understand the weaknesses of such systems: comparison to a classical…

Cited by 0SourceScholar
2016

Title assignment for automatic topic segments in TV broadcast news

ICASSP 2016accepted

This paper addresses the task of assigning a title to topic segments automatically extracted from TV Broadcast News video recordings. We propose to associate a topic segment with the title of a newspaper article collected on the web at the same date. The task implies pairing newspaper articles and t…

Cited by 0SourceScholar