2024
Interpretability Guarantees with Merlin-Arthur Classifiers
AISTATS 2024poster
We propose an interactive multi-agent classifier that provides provable interpretability guarantees even for complex agents such as neural networks. These guarantees consist of lower bounds on the mutual information between selected features and the classification decision. Our results are inspired…