← Search

Alina Maria Cristea

4 accepted papers

2024

Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages

COLING 2024main

Identifying the type of relationship between words (cognates, borrowings, inherited) provides a deeper insight into the history of a language and allows for a better characterization of language relatedness. In this paper, we propose a computational approach for discriminating between cognates and b…

2024

Verba volant, scripta volant? Don’t worry! There are computational solutions for protoword reconstruction

EMNLP 2024main

We introduce a new database of cognate words and etymons for the five main Romance languages, the most comprehensive one to date. We propose a strong benchmark for the automatic reconstruction of protowords for Romance languages, by applying a set of machine learning models and features on these dat…

2023

RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification

EMNLP 2023long main

The identification of cognates is a fundamental process in historical linguistics, on which any further research is based. Even though there are several cognate databases for Romance languages, they are rather scattered, incomplete, noisy, contain unreliable information, or have uncertain availabili…

Cited by 10SourceScholar
2021

Automatic Discrimination between Inherited and Borrowed Latin Words in Romance Languages

EMNLP 2021finding

In this paper, we address the problem of automatically discriminating between inherited and borrowed Latin words. We introduce a new dataset and investigate the case of Romance languages (Romanian, Italian, French, Spanish, Portuguese and Catalan), where words directly inherited from Latin coexist w…