Don’t Retrain, Just Rewrite: Countering Adversarial Perturbations by Rewriting Text
Can language models transform inputs to protect text classifiers against adversarial attacks? In this work, we present ATINTER, a model that intercepts and learns to rewrite adversarial inputs to make them non-adversarial for a downstream text classifier. Our experiments on four datasets and five at…