← Search

Emmanuel Morin

6 accepted papers

2025

AdminSet and AdminBERT: a Dataset and a Pre-trained Language Model to Explore the Unstructured Maze of French Administrative Documents

COLING 2025main

In recent years, Pre-trained Language Models(PLMs) have been widely used to analyze various documents, playing a crucial role in Natural Language Processing (NLP). However, administrative texts have rarely been used in information extraction tasks, even though this resource is available as open data…

2024

BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

ACL 2024findings

Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, offering potential applications across specialized domains such as healthcare and medicine. Despite the availability of various open-source LLMs tailored for health contexts, adapting general-purpose LLMs to the m…

2024

DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain

COLING 2024main

The biomedical domain has sparked a significant interest in the field of Natural Language Processing (NLP), which has seen substantial advancements with pre-trained language models (PLMs). However, comparing these models has proven challenging due to variations in evaluation protocols across differe…

2023

DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains

ACL 2023long

In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on general domain data, specialized ones have emerged to more effectively treat specific domains. In this paper, we propose…

Cited by 53SourcePDFScholar
2020

Data Selection for Bilingual Lexicon Induction from Specialized Comparable Corpora

COLING 2020main

Narrow specialized comparable corpora are often small in size. This particularity makes it difficult to build efficient models to acquire translation equivalents, especially for less frequent and rare words. One way to overcome this issue is to enrich the specialized corpora with out-of-domain resou…

Cited by 3SourcePDFScholar
2020

Error Analysis Applied to End-to-End Spoken Language Understanding

ICASSP 2020accepted

This paper presents a qualitative study of errors produced by an end-to-end spoken language understanding (SLU) system (speech signal to concepts) that reaches state of the art performance. Different studies are proposed to better understand the weaknesses of such systems: comparison to a classical…

Cited by 0SourceScholar