ACL 2024short4 citations

Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster

Agostina Calabrese, Leonardo Neves, Neil Shah, Maarten Bos, Björn Ross, Mirella Lapata, Francesco Barbieri

Abstract

Content moderators play a key role in keeping the conversation on social media healthy. While the high volume of content they need to judge represents a bottleneck to the moderation pipeline, no studies have explored how models could support them to make faster decisions. There is, by now, a vast body of research into detecting hate speech, sometimes explicitly motivated by a desire to help improve content moderation, but published research using real content moderators is scarce. In this work we investigate the effect of explanations on the speed of real-world moderators. Our experiments show that while generic explanations do not affect their speed and are often ignored, structured explanations lower moderators’ decision making time by 7.4%.

BibTeX
@inproceedings{calabrese-etal-2024-explainability,
    title = "Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster",
    author = {Calabrese, Agostina  and
      Neves, Leonardo  and
      Shah, Neil  and
      Bos, Maarten  and
      Ross, Bj{\"o}rn  and
      Lapata, Mirella  and
      Barbieri, Francesco},
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-short.38/",
    doi = "10.18653/v1/2024.acl-short.38",
    pages = "398--408"
}
Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster · ACL 2024