NeurIPS 2025poster0 citations

Manipulating Feature Visualizations with Gradient Slingshots

Dilyara Bareeva, Marina MC Höhne, Alexander Warnecke, Lukas Pirch, Klaus Robert Muller, Konrad Rieck, Sebastian Lapuschkin, Kirill Bykov

Abstract

Feature Visualization (FV) is a widely used technique for interpreting concepts learned by Deep Neural Networks (DNNs), which synthesizes input patterns that maximally activate a given feature. Despite its popularity, the trustworthiness of FV explanations has received limited attention. We introduce Gradient Slingshots, a novel method that enables FV manipulation without modifying model architecture or significantly degrading performance. By shaping new trajectories in off-distribution regions of a feature's activation landscape, we coerce the optimization process to converge to a predefined visualization. We evaluate our approach on several DNN architectures, demonstrating its ability to replace faithful FVs with arbitrary targets. These results expose a critical vulnerability: auditors relying solely on FV may accept entirely fabricated explanations. To mitigate this risk, we propose a straightforward defense and quantitatively demonstrate its effectiveness.

Explainable AIMechanistic InterpretabilityMachine LearningComputer Vision
BibTeX
@inproceedings{
bareeva2025manipulating,
title={Manipulating Feature Visualizations with Gradient Slingshots},
author={Dilyara Bareeva and Marina MC H{\"o}hne and Alexander Warnecke and Lukas Pirch and Klaus Robert Muller and Konrad Rieck and Sebastian Lapuschkin and Kirill Bykov},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=TgczQwE1Iu}
}