NeurIPS 2022accept73 citations

GlanceNets: Interpretable, Leak-proof Concept-based Models

Emanuele Marconato, Andrea Passerini, Stefano Teso

Abstract

There is growing interest in concept-based models (CBMs) that combine high-performance and interpretability by acquiring and reasoning with a vocabulary of high-level concepts. A key requirement is that the concepts be interpretable. Existing CBMs tackle this desideratum using a variety of heuristics based on unclear notions of interpretability, and fail to acquire concepts with the intended semantics. We address this by providing a clear definition of interpretability in terms of alignment between the model’s representation and an underlying data generation process, and introduce GlanceNets, a new CBM that exploits techniques from disentangled representation learning and open-set recognition to achieve alignment, thus improving the interpretability of the learned concepts. We show that GlanceNets, paired with concept-level supervision, achieve better alignment than state-of-the-art approaches while preventing spurious information from unintendedly leaking into the learned concepts.

explainabilityconcept-based modelsinterpretabilitydisentanglementconcept leakage
BibTeX
@inproceedings{
marconato2022glancenets,
title={GlanceNets: Interpretable, Leak-proof Concept-based Models},
author={Emanuele Marconato and Andrea Passerini and Stefano Teso},
booktitle={Advances in Neural Information Processing Systems},
editor={Alice H. Oh and Alekh Agarwal and Danielle Belgrave and Kyunghyun Cho},
year={2022},
url={https://openreview.net/forum?id=J7zY9j75GoG}
}