← Search

David Klindt

10 accepted papers

2026

Position: Causality is Key for Interpretability Claims to Generalise

ICML 2026poster

Interpretability research on large language models (LLMs) has produced methods that align model components to high-level concepts, yet their use has been accompanied by recurring failures: findings that do not generalise, and causal language that outruns the evidence. Our position is that Pearl’s ca…

Cited by 0SourceScholar
2025

AI-Generated Video Detection via Perceptual Straightening

NeurIPS 2025poster

The rapid advancement of generative AI enables highly realistic synthetic video, posing significant challenges for content authentication and raising urgent concerns about misuse. Existing detection methods often struggle with generalization and capturing subtle temporal inconsistencies. We propose…

Cited by 0SourceScholar
2025

Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders

ICML 2025poster

A recent line of work has shown promise in using sparse autoencoders (SAEs) to uncover interpretable features in neural network representations. However, the simple linear-nonlinear encoding mechanism in SAEs limits their ability to perform accurate sparse inference. Using compressed sensing theory,…

Cited by 0SourcePDFScholar
2025

Cross-Entropy Is All You Need To Invert the Data Generating Process

ICLR 2025oral

Supervised learning has become a cornerstone of modern machine learning, yet a comprehensive theory explaining its effectiveness remains elusive. Empirical phenomena, such as neural analogy-making and the linear representation hypothesis, suggest that supervised models can learn interpretable factor…

Cited by 2SourcePDFScholar
2025

Position: An Empirically Grounded Identifiability Theory Will Accelerate Self Supervised Learning Research

ICML 2025poster

Self-Supervised Learning (SSL) powers many current AI systems. As research interest and investment grow, the SSL design space continues to expand. The Platonic view of SSL, following the Platonic Representation Hypothesis (PRH), suggests that despite different methods and engineering approaches, all…

Cited by 0SourcePDFScholar
2024

Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning

NeurIPS 2024spotlight

While the impressive performance of modern neural networks is often attributed to their capacity to efficiently extract task-relevant features from data, the mechanisms underlying this *rich feature learning regime* remain elusive, with much of our theoretical understanding stemming from the opposin…

2024

Global Distortions from Local Rewards: Neural Coding Strategies in Path-Integrating Neural Systems

NeurIPS 2024poster

Grid cells in the mammalian brain are fundamental to spatial navigation, and therefore crucial to how animals perceive and interact with their environment. Traditionally, grid cells are thought support path integration through highly symmetric hexagonal lattice firing patterns. However, recent findi…

Cited by 0SourcePDFScholar
2024

Measuring Per-Unit Interpretability at Scale Without Humans

NeurIPS 2024poster

In today’s era, whatever we can measure at scale, we can optimize. So far, measuring the interpretability of units in deep neural networks (DNNs) for computer vision still requires direct human evaluation and is not scalable. As a result, the inner workings of DNNs remain a mystery despite the remar…

Cited by 1SourcePDFScholar
2020

System Identification with Biophysical Constraints: A Circuit Model of the Inner Retina

NeurIPS 2020spotlight

Visual processing in the retina has been studied in great detail at all levels such that a comprehensive picture of the retina's cell types and the many neural circuits they form is emerging. However, the currently best performing models of retinal function are black-box CNN models which are agnosti…

2017

Neural system identification for large populations separating “what” and “where”

NeurIPS 2017poster

Neuroscientists classify neurons into different types that perform similar computations at different locations in the visual field. Traditional methods for neural system identification do not capitalize on this separation of “what” and “where”. Learning deep convolutional feature spaces that are sh…