← Search

Marina MC Höhne

5 accepted papers

2025

Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework

NeurIPS 2025poster

Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large language models (LLMs) for natural language processing (NLP), current automated neuron-level feature description methods face…

Cited by 0SourceScholar
2025

Manipulating Feature Visualizations with Gradient Slingshots

NeurIPS 2025poster

Feature Visualization (FV) is a widely used technique for interpreting concepts learned by Deep Neural Networks (DNNs), which synthesizes input patterns that maximally activate a given feature. Despite its popularity, the trustworthiness of FV explanations has received limited attention. We introduc…

Cited by 0SourcecodeScholar
2024

CoSy: Evaluating Textual Explanations of Neurons

NeurIPS 2024poster

A crucial aspect of understanding the complex nature of Deep Neural Networks (DNNs) is the ability to explain learned concepts within their latent representations. While methods exist to connect neurons to human-understandable textual descriptions, evaluating the quality of these explanations is cha…

2023

Labeling Neural Representations with Inverse Recognition

NeurIPS 2023poster

Deep Neural Networks (DNNs) demonstrate remarkable capabilities in learning complex hierarchical data representations, but the nature of these representations remains largely unknown. Existing global explainability methods, such as Network Dissection, face limitations such as reliance on segmentatio…

2022

ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational Model

NeurIPS 2022accept

The need for interpretable models has fostered the development of self-explainable classifiers. Prior approaches are either based on multi-stage optimization schemes, impacting the predictive performance of the model, or produce explanations that are not transparent, trustworthy or do not capture th…