← Search

Frederik Pahde

2 accepted papers

2025

Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence

ICLR 2025poster

With a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable concepts in the latent space. Commonly, CAVs are computed by leveraging linear classifiers optimizing the *separability* o…

Cited by 5SourcePDFScholar
2024

From Hope to Safety: Unlearning Biases of Deep Models via Gradient Penalization in Latent Space

AAAI 2024technical

Deep Neural Networks are prone to learning spurious correlations embedded in the training data, leading to potentially biased predictions. This poses risks when deploying these models for high-stake decision-making, such as in medical applications. Current methods for post-hoc model correction eithe…