← Search

Richard Dufour

14 accepted papers

2025

A Benchmark of French ASR Systems Based on Error Severity

COLING 2025main

Automatic Speech Recognition (ASR) transcription errors are commonly assessed using metrics that compare them with a reference transcription, such as Word Error Rate (WER), which measures spelling deviations from the reference, or semantic score-based metrics. However, these approaches often overloo…

Cited by 0SourcePDFScholar
2025

ACL-rlg: A Dataset for Reading List Generation

COLING 2025main

Familiarizing oneself with a new scientific field and its existing literature can be daunting due to the large amount of available articles. Curated lists of academic references, or reading lists, compiled by experts, offer a structured way to gain a comprehensive overview of a domain or a specific…

2025

Identifying Reliable Evaluation Metrics for Scientific Text Revision

ACL 2025long

Evaluating text revision in scientific writing remains a challenge, as traditional metrics such as ROUGE and BERTScore primarily focus on similarity rather than capturing meaningful improvements. In this work, we analyse and identify the limitations of these metrics and explore alternative evaluatio…

2025

Persistent Homology of Topic Networks for the Prediction of Reader Curiosity

ACL 2025long

Reader curiosity, the drive to seek information, is crucial for textual engagement, yet remains relatively underexplored in NLP. Building on Loewenstein’s Information Gap Theory, we introduce a framework that models reader curiosity by quantifying semantic information gaps within a text’s semantic s…

2025

The Role of Natural Language Processing Tasks in Automatic Literary Character Network Construction

COLING 2025main

The automatic extraction of character networks from literary texts is generally carried out using natural language processing (NLP) cascading pipelines. While this approach is widespread, no study exists on the impact of low-level NLP tasks on their performance. In this article, we conduct such a st…

2024

A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks

COLING 2024main

The recent emergence of Large Language Models (LLMs) has enabled significant advances in the field of Natural Language Processing (NLP). While these new models have demonstrated superior performance on various tasks, their application and potential are still underexplored, both in terms of the diver…

Cited by 42SourcePDFScholar
2024

BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

ACL 2024findings

Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, offering potential applications across specialized domains such as healthcare and medicine. Despite the availability of various open-source LLMs tailored for health contexts, adapting general-purpose LLMs to the m…

2024

CASIMIR: A Corpus of Scientific Articles Enhanced with Multiple Author-Integrated Revisions

COLING 2024main

Writing a scientific article is a challenging task as it is a highly codified and specific genre, consequently proficiency in written communication is essential for effectively conveying research findings and ideas. In this article, we propose an original textual resource on the revision step of the…

Cited by 5SourcePDFScholar
2024

DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain

COLING 2024main

The biomedical domain has sparked a significant interest in the field of Natural Language Processing (NLP), which has seen substantial advancements with pre-trained language models (PLMs). However, comparing these models has proven challenging due to variations in evaluation protocols across differe…

2024

How Important Is Tokenization in French Medical Masked Language Models?

COLING 2024main

Subword tokenization has become the prevailing standard in the field of natural language processing (NLP) over recent years, primarily due to the widespread utilization of pre-trained language models. This shift began with Byte-Pair Encoding (BPE) and was later followed by the adoption of SentencePi…

2023

DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains

ACL 2023long

In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on general domain data, specialized ones have emerged to more effectively treat specific domains. In this paper, we propose…

Cited by 53SourcePDFScholar
2023

Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset

EMNLP 2023long main

While recent pre-trained transformer-based models can perform named entity recognition (NER) with great accuracy, their limited range remains an issue when applied to long documents such as whole novels. To alleviate this issue, a solution is to retrieve relevant context at the document level. Unfor…

Cited by 0SourcecodeScholar
2023

The Role of Global and Local Context in Named Entity Recognition

ACL 2023short

Pre-trained transformer-based models have recently shown great performance when applied to Named Entity Recognition (NER). As the complexity of their self-attention mechanism prevents them from processing long documents at once, these models are usually applied in a sequential fashion. Such an appro…

2019

Similarity Metric Based on Siamese Neural Networks for Voice Casting

ICASSP 2019accepted

Dubbing contributes to a larger international distribution of multimedia documents. It aims to replace the original voice in a source language by a new one in a target language. For now, the target voice selection procedure, called voice casting, is manually performed by human experts. This selectio…

Cited by 0SourceScholar