Cognate Detection for Historical Language Reconstruction of Proto-Sabean Languages: the Case of Ge’ez, Tigrinya, and Amharic
Elleni Sisay Temesgen, Hellina Hailu Nigatu, Fitsum Assamnew Andargie
Abstract
As languages evolve, we risk losing ancestral languages. In this paper, we explore Historical Language Reconstruction (HLR) for Proto-Sabean languages, starting with the identification of cognates–sets of words in different related languages that are derived from the same ancestral language. We (1) collect semantically related words in three Afro-Semitic languages from a three-way dictionary (2) work with linguists to identify cognates and reconstruct the proto-form of the cognates, (3) experiment with three automatic cognate detection methods and extract cognates from the semantically related words. We then experiment with in-context learning with GPT-4o to generate the proto-language from the cognates and use Sequence-to-Sequence (Seq2Seq) models for HLR.
BibTeX
@inproceedings{temesgen-etal-2025-cognate,
title = "Cognate Detection for Historical Language Reconstruction of Proto-Sabean Languages: the Case of {G}e{'}ez, {T}igrinya, and {A}mharic",
author = "Temesgen, Elleni Sisay and
Nigatu, Hellina Hailu and
Andargie, Fitsum Assamnew",
editor = "Rambow, Owen and
Wanner, Leo and
Apidianaki, Marianna and
Al-Khalifa, Hend and
Eugenio, Barbara Di and
Schockaert, Steven",
booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
month = jan,
year = "2025",
address = "Abu Dhabi, UAE",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.coling-main.496/",
pages = "7415--7422"
}