COLING 2025main0 citations

Semi-Automated Construction of Sense-Annotated Datasets for Practically Any Language

Jai Riley, Bradley M. Hauer, Nafisa Sadaf Hriti, Guoqing Luo, Amir Reza Mirzaei, Ali Rafiei, Hadi Sheikhi, Mahvash Siavashpour

Abstract

High-quality sense-annotated datasets are vital for evaluating and comparing WSD systems. We present a novel approach to creating parallel sense-annotated datasets, which can be applied to any language that English can be translated into. The method incorporates machine translation, word alignment, sense projection, and sense filtering to produce silver annotations, which can then be revised manually to obtain gold datasets. By applying our method to Farsi, Chinese, and Bengali, we produce new parallel benchmark datasets, which are vetted by native speakers of each language. Our automatically-generated silver datasets are of higher quality than the annotations obtained with recent multilingual WSD systems, particularly on non-European languages.

BibTeX
@inproceedings{riley-etal-2025-semi,
    title = "Semi-Automated Construction of Sense-Annotated Datasets for Practically Any Language",
    author = "Riley, Jai  and
      Hauer, Bradley M.  and
      Hriti, Nafisa Sadaf  and
      Luo, Guoqing  and
      Mirzaei, Amir Reza  and
      Rafiei, Ali  and
      Sheikhi, Hadi  and
      Siavashpour, Mahvash  and
      Tavakoli, Mohammad  and
      Shi, Ning  and
      Kondrak, Grzegorz",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.419/",
    pages = "6270--6284"
}
Semi-Automated Construction of Sense-Annotated Datasets for Practically Any Language · COLING 2025