← Search

Tomaž Erjavec

4 accepted papers

2024

Gos 2: A New Reference Corpus of Spoken Slovenian

COLING 2024main

This paper introduces a new version of the Gos reference corpus of spoken Slovenian, which was recently extended to more than double the original size (300 hours, 2.4 million words) by adding speech recordings and transcriptions from two related initiatives, the Gos VideoLectures corpus of public ac…

2024

SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene

COLING 2024main

This paper introduces the upgrade of a training corpus for linguistic annotation of modern standard Slovene. The enhancement spans both the size of the corpus and the depth of annotation layers. The revised SUK 1.0 corpus, building on its predecessor ssj500k 2.3, has doubled in size, containing over…

2024

Towards an Ideal Tool for Learner Error Annotation

COLING 2024main

Annotation and analysis of corrections in learner corpora have always presented technical challenges, mainly on account of the fact that until now there has not been any standard tool available, and that original and corrected versions of texts have been mostly stored together rather than treated as…

Cited by 4SourcePDFScholar
2022

Dealing with Abbreviations in the Slovenian Biographical Lexicon

EMNLP 2022main

Abbreviations present a significant challenge for NLP systems because they cause tokenization and out-of-vocabulary errors. They can also make the text less readable, especially in reference printed books, where they are extensively used. Abbreviations are especially problematic in low-resource sett…