2021
Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP Models
ACL 2021long
In this paper, we introduce Integrated Directional Gradients (IDG), a method for attributing importance scores to groups of features, indicating their relevance to the output of a neural network model for a given input. The success of Deep Neural Networks has been attributed to their ability to capt…