Overlooked Factors in Concept-Based Explanations: Dataset Choice, Concept Learnability, and Human Capability
Concept-based interpretability methods aim to explain a deep neural network model's components and predictions using a pre-defined set of semantic concepts. These methods evaluate a trained model on a new, "probe" dataset and correlate the model's outputs with concepts labeled in that dataset. Despi…