NAACL 2024long16 citations

Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark

Stephen Mayhew, Terra Blevins, Shuheng Liu, Marek Šuppa, Hila Gonen, Joseph Marvin Imperial, Börje F. Karlsson, Peiqin Lin

Abstract

We introduce Universal NER (UNER), an open, community-driven project to develop gold-standard NER benchmarks in many languages. The overarching goal of UNER is to provide high-quality, cross-lingually consistent annotations to facilitate and standardize multilingual NER research. UNER v1 contains 19 datasets annotated with named entities in a cross-lingual consistent schema across 13 diverse languages. In this paper, we detail the dataset creation and composition of UNER; we also provide initial modeling baselines on both in-language and cross-lingual learning settings. We will release the data, code, and fitted models to the public.

BibTeX
@inproceedings{mayhew-etal-2024-universal,
    title = "Universal {NER}: A Gold-Standard Multilingual Named Entity Recognition Benchmark",
    author = {Mayhew, Stephen  and
      Blevins, Terra  and
      Liu, Shuheng  and
      {\v{S}}uppa, Marek  and
      Gonen, Hila  and
      Imperial, Joseph Marvin  and
      Karlsson, B{\"o}rje F.  and
      Lin, Peiqin  and
      Ljube{\v{s}}i{\'c}, Nikola  and
      Miranda, LJ  and
      Plank, Barbara  and
      Riabi, Arij  and
      Pinter, Yuval},
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.243/",
    doi = "10.18653/v1/2024.naacl-long.243",
    pages = "4322--4337"
}
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark · NAACL 2024