← Search

Mihaela-Claudia Cercel

4 accepted papers

2025

GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering

ACL 2025finding

Pre-trained language models have shown remarkable performance in recent years, setting a new paradigm for natural language processing (NLP) research. The legal domain has received some attention from the NLP community, in part due to its textual nature. Question answering (QA) systems represent some…

2025

MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language

EMNLP 2025

This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced between the Romanian language spoken in Romania and the Republic of Moldova. We further propose a multi-target adversarial

Cited by 0SourcePDFScholar
2025

RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation

COLING 2025main

Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language processing, most of the datasets available for summarization are in English, posing challenges for developing summarization mo…

2024

Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain Setups

EMNLP 2024main

Complex Word Identification (CWI) is an essential step in the lexical simplification task and has recently become a task on its own. Some variations of this binary classification task have emerged, such as lexical complexity prediction (LCP) and complexity evaluation of multi-word expressions (MWE).…