← Search

Sourav Bhabesh

1 accepted papers

2023

Towards Building a Robust Toxicity Predictor

ACL 2023industry

Recent NLP literature pays little attention to the robustness of toxicity language predictors, while these systems are most likely to be used in adversarial contexts. This paper presents a novel adversarial attack, \texttt{ToxicTrap}, introducing small word-level perturbations to fool SOTA text clas…