EMNLP 2021system demonstrations7 citations

LexiClean: An annotation tool for rapid multi-task lexical normalisation

Tyler Bikaun, Tim French, Melinda Hodkiewicz, Michael Stewart, Wei Liu

Abstract

NLP systems are often challenged by difficulties arising from noisy, non-standard, and domain specific corpora. The task of lexical normalisation aims to standardise such corpora, but currently lacks suitable tools to acquire high-quality annotated data to support deep learning based approaches. In this paper, we present LexiClean, the first open-source web-based annotation tool for multi-task lexical normalisation. LexiClean’s main contribution is support for simultaneous in situ token-level modification and annotation that can be rapidly applied corpus wide. We demonstrate the usefulness of our tool through a case study on two sets of noisy corpora derived from the specialised-domain of industrial mining. We show that LexiClean allows for the rapid and efficient development of high-quality parallel corpora. A demo of our system is available at: https://youtu.be/P7_ooKrQPDU.

BibTeX
@inproceedings{bikaun-etal-2021-lexiclean,
    title = "{L}exi{C}lean: An annotation tool for rapid multi-task lexical normalisation",
    author = "Bikaun, Tyler  and
      French, Tim  and
      Hodkiewicz, Melinda  and
      Stewart, Michael  and
      Liu, Wei",
    editor = "Adel, Heike  and
      Shi, Shuming",
    booktitle = "Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations",
    month = nov,
    year = "2021",
    address = "Online and Punta Cana, Dominican Republic",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.emnlp-demo.25/",
    doi = "10.18653/v1/2021.emnlp-demo.25",
    pages = "212--219"
}