2020
Generative causal explanations of black-box classifiers
NeurIPS 2020poster
We develop a method for generating causal post-hoc explanations of black-box classifiers based on a learned low-dimensional representation of the data. The explanation is causal in the sense that changing learned latent factors produces a change in the classifier output statistics. To construct thes…