← Search

Jaka Čibej

2 accepted papers

2024

SI-NLI: A Slovene Natural Language Inference Dataset and Its Evaluation

COLING 2024main

Natural language inference (NLI) is an important language understanding benchmark. Two deficiencies of this benchmark are: i) most existing NLI datasets exist for English and a few other well-resourced languages, and ii) most NLI datasets are formed with a narrow set of annotators’ instructions, all…

2024

SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene

COLING 2024main

This paper introduces the upgrade of a training corpus for linguistic annotation of modern standard Slovene. The enhancement spans both the size of the corpus and the depth of annotation layers. The revised SUK 1.0 corpus, building on its predecessor ssj500k 2.3, has doubled in size, containing over…