← Search

Stephan Wäldchen

3 accepted papers

2024

Interpretability Guarantees with Merlin-Arthur Classifiers

AISTATS 2024poster

We propose an interactive multi-agent classifier that provides provable interpretability guarantees even for complex agents such as neural networks. These guarantees consist of lower bounds on the mutual information between selected features and the classification decision. Our results are inspired…

2022

Training Characteristic Functions with Reinforcement Learning: XAI-methods play Connect Four

ICML 2022oral

Characteristic functions (from cooperative game theory) are able to evaluate partial inputs and form the basis for attribution methods like Shapley values. These attribution methods allow us to measure how important each input component is for the function output—one of the goals of explainable AI (…