EMNLP 2021finding6 citations

Do UD Trees Match Mention Spans in Coreference Annotations?

Martin Popel, Zdeněk Žabokrtský, Anna Nedoluzhko, Michal Novák, Daniel Zeman

Abstract

One can find dozens of data resources for various languages in which coreference - a relation between two or more expressions that refer to the same real-world entity - is manually annotated. One could also assume that such expressions usually constitute syntactically meaningful units; however, mention spans have been annotated simply by delimiting token intervals in most coreference projects, i.e., independently of any syntactic representation. We argue that it could be advantageous to make syntactic and coreference annotations convergent in the long term. We present a pilot empirical study focused on matches and mismatches between hand-annotated linear mention spans and automatically parsed syntactic trees that follow Universal Dependencies conventions. The study covers 9 datasets for 8 different languages.

BibTeX
@inproceedings{popel-etal-2021-ud-trees,
    title = "Do {UD} Trees Match Mention Spans in Coreference Annotations?",
    author = "Popel, Martin  and
      {\v{Z}}abokrtsk{\'y}, Zden{\v{e}}k  and
      Nedoluzhko, Anna  and
      Nov{\'a}k, Michal  and
      Zeman, Daniel",
    editor = "Moens, Marie-Francine  and
      Huang, Xuanjing  and
      Specia, Lucia  and
      Yih, Scott Wen-tau",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
    month = nov,
    year = "2021",
    address = "Punta Cana, Dominican Republic",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.findings-emnlp.303/",
    doi = "10.18653/v1/2021.findings-emnlp.303",
    pages = "3570--3576"
}