← Search

Laurent Besacier

22 accepted papers

2025

ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models

COLING 2025main

Research on Large Language Models (LLMs) has recently witnessed an increasing interest in extending the models’ context size to better capture dependencies within long documents. While benchmarks have been proposed to assess long-range abilities, existing efforts primarily considered generic tasks t…

2025

Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection

COLING 2025main

While crowdsourcing is an established solution for facilitating and scaling the collection of speech data, the involvement of non-experts necessitates protocols to ensure final data quality. To reduce the costs of these essential controls, this paper investigates the use of Speech Foundation Models…

2022

Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation Models

ACL 2022long

Multi-encoder models are a broad family of context-aware neural machine translation systems that aim to improve translation quality by encoding document-level contextual information alongside the current sentence. The context encoding is undertaken by contextual parameters, trained on document-level…

2022

Fashioning Local Designs from Generic Speech Technologies in an Australian Aboriginal Community

COLING 2022main

An increasing number of papers have been addressing issues related to low-resource languages and the transcription bottleneck paradigm. After several years spent in Northern Australia, where some of the strongest Aboriginal languages are spoken, we could observe a gap between the motivations depicte…

Cited by 7SourcePDFScholar
2022

Learning From Failure: Data Capture in an Australian Aboriginal Community

ACL 2022long

Most low resource language technology development is premised on the need to collect data for training statistical models. When we follow the typical process of recording and transcribing text for small Indigenous languages, we hit up against the so-called “transcription bottleneck.” Therefore it is…

Cited by 13SourcePDFScholar
2022

SMaLL-100: Introducing Shallow Multilingual Machine Translation Model for Low-Resource Languages

EMNLP 2022main

In recent years, multilingual machine translation models have achieved promising performance on low-resource language pairs by sharing information between similar languages, thus enabling zero-shot translation. To overcome the “curse of multilinguality”, these models often opt for scaling up the num…

2022

Weakly Supervised Word Segmentation for Computational Language Documentation

ACL 2022long

Word and morpheme segmentation are fundamental steps of language documentation as they allow to discover lexical units in a language for which the lexicon is unknown. However, in most language documentation scenarios, linguists do not start from a blank page: they may already have a pre-existing dic…

2022

What Do Compressed Multilingual Machine Translation Models Forget?

EMNLP 2022finding

Recently, very large pre-trained models achieve state-of-the-art results in various natural language processing (NLP) tasks, but their size makes it more challenging to apply them in resource-constrained environments. Compression techniques allow to drastically reduce the size of the models and ther…

2021

An Empirical Study of End-To-End Simultaneous Speech Translation Decoding Strategies

ICASSP 2021accepted

This paper proposes a decoding strategy for end-to-end simultaneous speech translation. We leverage end-to-end models trained in offline mode and conduct an empirical study for two language pairs (English-to-German and English-to-Portuguese). We also investigate different output token granularities…

Cited by 0SourceScholar
2021

Lightweight Adapter Tuning for Multilingual Speech Translation

ACL 2021short

Adapter modules were recently introduced as an efficient alternative to fine-tuning in NLP. Adapter tuning consists in freezing pre-trained parameters of a model and injecting lightweight modules between layers, resulting in the addition of only a small number of task-specific trainable parameters.…

2021

Multilingual Unsupervised Neural Machine Translation with Denoising Adapters

EMNLP 2021main

We consider the problem of multilingual unsupervised machine translation, translating to and from languages that only have monolingual data by using auxiliary parallel language pairs. For this problem the standard procedure so far to leverage the monolingual data is _back-translation_, which is comp…

2021

Task Agnostic and Task Specific Self-Supervised Learning from Speech with LeBenchmark

NeurIPS 2021poster

Self-Supervised Learning (SSL) has yielded remarkable improvements in many different domains including computer vision, natural language processing and speech processing by leveraging large amounts of unlabeled data. In the specific context of speech, however, and despite promising results, there ex…

Cited by 41SourceScholar
2020

A Data Efficient End-to-End Spoken Language Understanding Architecture

ICASSP 2020accepted

End-to-end architectures have been recently proposed for spoken language understanding (SLU) and semantic parsing. Based on a large amount of data, those models learn jointly acoustic and linguistic-sequential features. Such architectures give very good results in the context of domain, intent and s…

Cited by 0SourceScholar
2020

Dual-decoder Transformer for Joint Automatic Speech Recognition and Multilingual Speech Translation

COLING 2020main

We introduce dual-decoder Transformer, a new model architecture that jointly performs automatic speech recognition (ASR) and multilingual speech translation (ST). Our models are based on the original Transformer architecture (Vaswani et al., 2017) but consist of two decoders, each responsible for on…

2020

Enabling Interactive Transcription in an Indigenous Community

COLING 2020main

We propose a novel transcription workflow which combines spoken term detection and human-in-the-loop, together with a pilot experiment. This work is grounded in an almost zero-resource scenario where only a few terms have so far been identified, involving two endangered languages. We show that in th…

2020

Online Versus Offline NMT Quality: An In-depth Analysis on English-German and German-English

COLING 2020main

We conduct in this work an evaluation study comparing offline and online neural machine translation architectures. Two sequence-to-sequence models: convolutional Pervasive Attention (Elbayad et al. 2018) and attention-based Transformer (Vaswani et al. 2017) are considered. We investigate, for both a…

2019

Models of Visually Grounded Speech Signal Pay Attention to Nouns: A Bilingual Experiment on English and Japanese

ICASSP 2019accepted

We investigate the behaviour of attention in neural models of visually grounded speech trained on two languages: English and Japanese. Experimental results show that attention focuses on nouns and this behaviour holds true for two very typologically different languages. We also draw parallels betwee…

Cited by 0SourceScholar
2018

ASR Performance Prediction on Unseen Broadcast Programs Using Convolutional Neural Networks

ICASSP 2018accepted

In this paper, we address a relatively new task: prediction of ASR performance on unseen broadcast programs. We first propose an heterogenous French corpus dedicated to this task. Two prediction approaches are compared: a state-of-the-art performance prediction based on regression (engineered featur…

Cited by 0SourceScholar
2018

Bayesian Models for Unit Discovery on a Very Low Resource Language

ICASSP 2018accepted

Developing speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to u…

Cited by 0SourceScholar
2018

End-to-End Automatic Speech Translation of Audiobooks

ICASSP 2018accepted

We investigate end-to-end speech-to-text translation on a corpus of audiobooks specifically augmented for this task. Previous works investigated the extreme case where source language transcription is not available during learning nor decoding, but we also study a midway case where source language t…

Cited by 0SourceScholar
2018

Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop

ICASSP 2018accepted

We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translate…

Cited by 0SourceScholar
2016

OCR-aided person annotation and label propagation for speaker modeling in TV shows

ICASSP 2016accepted

In this paper, we present an approach for minimizing human effort in manual speaker annotation. Label propagation is used at each iteration of an active learning cycle. More precisely, a selection strategy for choosing the most suitable speech track to be labeled is proposed. Four different selectio…

Cited by 0SourceScholar