NAACL 2021long53 citations

Discourse Probing of Pretrained Language Models

Fajri Koto, Jey Han Lau, Timothy Baldwin

Abstract

Existing work on probing of pretrained language models (LMs) has predominantly focused on sentence-level syntactic tasks. In this paper, we introduce document-level discourse probing to evaluate the ability of pretrained LMs to capture document-level relations. We experiment with 7 pretrained LMs, 4 languages, and 7 discourse probing tasks, and find BART to be overall the best model at capturing discourse — but only in its encoder, with BERT performing surprisingly well as the baseline model. Across the different models, there are substantial differences in which layers best capture discourse information, and large disparities between models.

BibTeX
@inproceedings{koto-etal-2021-discourse,
    title = "Discourse Probing of Pretrained Language Models",
    author = "Koto, Fajri  and
      Lau, Jey Han  and
      Baldwin, Timothy",
    editor = "Toutanova, Kristina  and
      Rumshisky, Anna  and
      Zettlemoyer, Luke  and
      Hakkani-Tur, Dilek  and
      Beltagy, Iz  and
      Bethard, Steven  and
      Cotterell, Ryan  and
      Chakraborty, Tanmoy  and
      Zhou, Yichao",
    booktitle = "Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jun,
    year = "2021",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.naacl-main.301/",
    doi = "10.18653/v1/2021.naacl-main.301",
    pages = "3849--3864"
}
Discourse Probing of Pretrained Language Models · NAACL 2021