ACL 2021long42 citations

Reliability Testing for Natural Language Processing Systems

Samson Tan, Shafiq Joty, Kathy Baxter, Araz Taeihagh, Gregory A. Bennett, Min-Yen Kan

Abstract

Questions of fairness, robustness, and transparency are paramount to address before deploying NLP systems. Central to these concerns is the question of reliability: Can NLP systems reliably treat different demographics fairly and function correctly in diverse and noisy environments? To address this, we argue for the need for reliability testing and contextualize it among existing work on improving accountability. We show how adversarial attacks can be reframed for this goal, via a framework for developing reliability tests. We argue that reliability testing — with an emphasis on interdisciplinary collaboration — will enable rigorous and targeted testing, and aid in the enactment and enforcement of industry standards.

BibTeX
@inproceedings{tan-etal-2021-reliability,
    title = "Reliability Testing for Natural Language Processing Systems",
    author = "Tan, Samson  and
      Joty, Shafiq  and
      Baxter, Kathy  and
      Taeihagh, Araz  and
      Bennett, Gregory A.  and
      Kan, Min-Yen",
    editor = "Zong, Chengqing  and
      Xia, Fei  and
      Li, Wenjie  and
      Navigli, Roberto",
    booktitle = "Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)",
    month = aug,
    year = "2021",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.acl-long.321/",
    doi = "10.18653/v1/2021.acl-long.321",
    pages = "4153--4169"
}
Reliability Testing for Natural Language Processing Systems · ACL 2021