Enhancing Interpretability for Vision Models via Shapley Value Optimization
Kanglong Fan, Yunqiao Yang, Chen Ma
Abstract
Deep neural networks have demonstrated remarkable performance across various domains, yet their decision-making processes remain opaque. Although many explanation methods are dedicated to bringing the obscurity of DNNs to light, they exhibit significant limitations: post-hoc explanation methods often struggle to faithfully reflect model behaviors, while self-explaining neural networks sacrifice performance and compatibility due to their specialized architectural designs. To address these challenges, we propose a novel self-explaining framework that integrates Shapley value estimation as an auxiliary task during training, which achieves two key advancements: 1) a fair allocation of the model prediction scores to image patches, ensuring explanations inherently align with the model
BibTeX
@inproceedings{aaai2026_enhancinginterpr,
title = {Enhancing Interpretability for Vision Models via Shapley Value Optimization},
author = {Kanglong Fan and Yunqiao Yang and Chen Ma},
booktitle = {AAAI 2026},
year = {2026}
}