← Search

Mickael Rouvier

9 accepted papers

2025

A Benchmark of French ASR Systems Based on Error Severity

COLING 2025main

Automatic Speech Recognition (ASR) transcription errors are commonly assessed using metrics that compare them with a reference transcription, such as Word Error Rate (WER), which measures spelling deviations from the reference, or semantic score-based metrics. However, these approaches often overloo…

Cited by 0SourcePDFScholar
2024

A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks

COLING 2024main

The recent emergence of Large Language Models (LLMs) has enabled significant advances in the field of Natural Language Processing (NLP). While these new models have demonstrated superior performance on various tasks, their application and potential are still underexplored, both in terms of the diver…

Cited by 42SourcePDFScholar
2024

BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

ACL 2024findings

Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, offering potential applications across specialized domains such as healthcare and medicine. Despite the availability of various open-source LLMs tailored for health contexts, adapting general-purpose LLMs to the m…

2024

DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain

COLING 2024main

The biomedical domain has sparked a significant interest in the field of Natural Language Processing (NLP), which has seen substantial advancements with pre-trained language models (PLMs). However, comparing these models has proven challenging due to variations in evaluation protocols across differe…

2024

How Important Is Tokenization in French Medical Masked Language Models?

COLING 2024main

Subword tokenization has become the prevailing standard in the field of natural language processing (NLP) over recent years, primarily due to the widespread utilization of pre-trained language models. This shift began with Byte-Pair Encoding (BPE) and was later followed by the adoption of SentencePi…

2024

Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset

ICASSP 2024accepted

The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech coll…

Cited by 0SourceScholar
2023

DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains

ACL 2023long

In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on general domain data, specialized ones have emerged to more effectively treat specific domains. In this paper, we propose…

Cited by 53SourcePDFScholar
2023

Jeffreys Divergence-Based Regularization of Neural Network Output Distribution Applied to Speaker Recognition

ICASSP 2023accepted

A new loss function for speaker recognition with deep neural network is proposed, based on Jeffreys Divergence. Adding this divergence to the cross-entropy loss function allows to maximize the target value of the output distribution while smoothing the non-target values. This objective function prov…

Cited by 0SourceScholar