← Search

Kathleen Fraser

2 accepted papers

2022

Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation Vectors

ACL 2022long

Robustness of machine learning models on ever-changing real-world data is critical, especially for applications affecting human well-being such as content moderation. New kinds of abusive language continually emerge in online discussions in response to current events (e.g., COVID-19), and the deploy…

2022

Necessity and Sufficiency for Explaining Text Classifiers: A Case Study in Hate Speech Detection

NAACL 2022long

We present a novel feature attribution method for explaining text classifiers, and analyze it in the context of hate speech detection. Although feature attribution models usually provide a single importance score for each token, we instead provide two complementary and theoretically-grounded scores…