NeurIPS 2023poster6 citations

Concept Distillation: Leveraging Human-Centered Explanations for Model Improvement

Avani Gupta, Saurabh Saini, P J Narayanan

Abstract

Humans use abstract *concepts* for understanding instead of hard features. Recent interpretability research has focused on human-centered concept explanations of neural networks. Concept Activation Vectors (CAVs) estimate a model's sensitivity and possible biases to a given concept. We extend CAVs from post-hoc analysis to ante-hoc training to reduce model bias through fine-tuning using an additional *Concept Loss*. Concepts are defined on the final layer of the network in the past. We generalize it to intermediate layers, including the last convolution layer. We also introduce *Concept Distillation*, a method to define rich and effective concepts using a pre-trained knowledgeable model as the teacher. Our method can sensitize or desensitize a model towards concepts. We show applications of concept-sensitive training to debias several classification problems. We also show a way to induce prior knowledge into a reconstruction problem. We show that concept-sensitive training can improve model interpretability, reduce biases, and induce prior knowledge.

Human Centered ConceptsML interpretabilityXAI based Model ImprovementDebiasing
BibTeX
@inproceedings{
gupta2023concept,
title={Concept Distillation: Leveraging Human-Centered Explanations for Model Improvement},
author={Avani Gupta and Saurabh Saini and P J Narayanan},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems},
year={2023},
url={https://openreview.net/forum?id=arkmhtYLL6}
}
Concept Distillation: Leveraging Human-Centered Explanations for Model Improvement · NeurIPS 2023