2022
Input-specific Attention Subnetworks for Adversarial Detection
ACL 2022findings
Self-attention heads are characteristic of Transformer models and have been well studied for interpretability and pruning. In this work, we demonstrate an altogether different utility of attention heads, namely for adversarial detection. Specifically, we propose a method to construct input-specific…