← Search

Thomas Fel

30 accepted papers

2026

Block Recurrent Dynamics in Vision Transformers

ICLR 2026poster

As Vision Transformers (ViTs) become standard backbones across vision, a mechanistic account of their computational phenomenology is now essential. Despite architectural cues that hint at dynamical structure, there is no settled framework that interprets Transformer depth as a well-characterized flo…

Cited by 0SourcecodeScholar
2026

Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings

ICLR 2026poster

Vision–language models (VLMs) align images and text with remarkable success, yet the geometry of their shared embedding space remains poorly understood. To probe this geometry, we begin from the Iso-Energy Assumption, which exploits cross-modal redundancy: a concept that is truly shared should exhi…

Cited by 0SourceScholar
2026

Interpreting Physics in Video World Models

ICML 2026poster

A long-standing question in physical reasoning is whether video-based models need to rely on factorized representations of physical variables in order to make physically accurate predictions, or whether they can implicitly represent such variables in a distributed manner. While modern video world mo…

Cited by 0SourceScholar
2026

Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry

ICLR 2026poster

DINOv2 sees the world well enough to guide robots and segment images, but we still do not know what it sees. We conduct the first comprehensive analysis of DINOv2’s representational structure using overcomplete dictionary learning, extracting over 32,000 visual concepts in what constitutes the large…

Cited by 0SourceScholar
2026

Priors in time: Missing inductive biases for language model interpretability

ICLR 2026poster

A central aim of interpretability tools applied to language models is to recover meaningful concepts from model activations. Existing feature extraction methods focus on single activations regardless of the context, implicitly assuming independence (and therefore stationarity). This leaves open whet…

Cited by 0SourcecodeScholar
2026

Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders

ICLR 2026poster

Despite their impressive performance, generative image models trained on large-scale datasets frequently fail to produce images with seemingly simple concepts -- e.g., human hands or objects appearing in groups of four -- that are reasonably expected to appear in the training data. These failure mod…

Cited by 0SourceScholar
2025

An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN Architectures

ICML 2025poster

Orthogonal convolutional layers are valuable components in multiple areas of machine learning, such as adversarial robustness, normalizing flows, GANs, and Lipschitz-constrained models. Their ability to preserve norms and ensure stable gradient propagation makes them valuable for a large range of pr…

2025

Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models

ICML 2025poster

Sparse Autoencoders (SAEs) have emerged as a powerful framework for machine learning interpretability, enabling the unsupervised decomposition of model representations into a dictionary of abstract, human-interpretable concepts. However, we reveal a fundamental limitation: SAEs exhibit severe instab…

Cited by 2SourcePDFScholar
2025

From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit

NeurIPS 2025poster

Motivated by the hypothesis that neural network representations encode abstract, interpretable features as linearly accessible, approximately orthogonal directions, sparse autoencoders (SAEs) have become a popular tool in interpretability literature. However, recent work has demonstrated phenomenolo…

Cited by 0SourceScholar
2025

One Wave To Explain Them All: A Unifying Perspective On Feature Attribution

ICML 2025poster

Feature attribution methods aim to improve the transparency of deep neural networks by identifying the input features that influence a model's decision. Pixel-based heatmaps have become the standard for attributing features to high-dimensional inputs, such as images, audio representations, and volum…

2025

Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry

NeurIPS 2025poster

Sparse Autoencoders (SAEs) are widely used to interpret neural networks by identifying meaningful concepts from their representations. However, do SAEs truly uncover all concepts a model relies on, or are they inherently biased toward certain kinds of concepts? We introduce a unified framework that…

Cited by 0SourceScholar
2025

Unearthing Skill-level Insights for Understanding Trade-offs of Foundation Models

ICLR 2025poster

With models getting stronger, evaluations have grown more complex, testing multiple skills in one benchmark and even in the same instance at once. However, skill-wise performance is obscured when inspecting aggregate accuracy, under-utilizing the rich signal modern benchmarks contain. We propose an…

Cited by 2SourcePDFScholar
2025

Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment

ICML 2025poster

We present Universal Sparse Autoencoders (USAEs), a framework for uncovering and aligning interpretable concepts spanning multiple pretrained deep neural networks. Unlike existing concept-based interpretability methods, which focus on a single model, USAEs jointly learn a universal concept space tha…

Cited by 4SourcePDFScholar
2025

Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models

NeurIPS 2025poster

Humans are able to recognize objects based on both local texture cues and the configuration of object parts, yet contemporary vision models primarily harvest local texture cues, yielding brittle, non-compositional features. Work on shape-vs-texture bias has pitted shape and texture representations i…

Cited by 0SourceScholar
2024

Latent Representation Matters: Human-like Sketches in One-shot Drawing Tasks

NeurIPS 2024poster

Humans can effortlessly draw new categories from a single exemplar, a feat that has long posed a challenge for generative models. However, this gap has started to close with recent advances in diffusion models. This one-shot drawing task requires powerful inductive biases that have not been systemat…

Cited by 0SourcePDFScholar
2024

Saliency strikes back: How filtering out high frequencies improves white-box explanations

ICML 2024poster

Attribution methods correspond to a class of explainability methods (XAI) that aim to assess how individual inputs contribute to a model's decision-making process. We have identified a significant limitation in one type of attribution methods, known as ``white-box" methods. Although highly efficient…

Cited by 4SourcePDFScholar
2024

Understanding Visual Feature Reliance through the Lens of Complexity

NeurIPS 2024poster

Recent studies suggest that deep learning models' inductive bias towards favoring simpler features may be an origin of shortcut learning. Yet, there has been limited focus on understanding the complexities of the myriad features that models learn. In this work, we introduce a new metric for quantify…

Cited by 5SourcePDFScholar
2023

A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation

NeurIPS 2023spotlight

In recent years, concept-based approaches have emerged as some of the most promising explainability methods to help us interpret the decisions of Artificial Neural Networks (ANNs). These methods seek to discover intelligible visual ``concepts'' buried within the complex patterns of ANN activations i…

Cited by 56SourcePDFScholar
2023

COCKATIEL: COntinuous Concept ranKed ATtribution with Interpretable ELements for explaining neural net classifiers on NLP

ACL 2023findings

Transformer architectures are complex and their use in NLP, while it has engendered many successes, makes their interpretability or explainability challenging. Recent debates have shown that attention maps and attribution methods are unreliable (Pruthi et al., 2019; Brunner et al., 2019). In this pa…

2023

CRAFT: Concept Recursive Activation FacTorization for Explainability

CVPR 2023poster

Attribution methods are a popular class of explainability methods that use heatmaps to depict the most important areas of an image that drive a model decision. Nevertheless, recent work has shown that these methods have limited utility in practice, presumably because they only highlight the most sal…

2023

Diffusion Models as Artists: Are we Closing the Gap between Humans and Machines?

ICML 2023oral

An important milestone for AI is the development of algorithms that can produce drawings that are indistinguishable from those of humans. Here, we adapt the ''diversity vs. recognizability'' scoring framework from Boutin et al (2022) and find that one-shot diffusion models have indeed started to clo…

2023

Don't Lie to Me! Robust and Efficient Explainability With Verified Perturbation Analysis

CVPR 2023poster

A variety of methods have been proposed to try to explain how deep neural networks make their decisions. Key to those approaches is the need to sample the pixel space efficiently in order to derive importance maps. However, it has been shown that the sampling methods used to date introduce biases an…

Cited by 42SourcePDFScholar
2023

On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport Perspective

NeurIPS 2023poster

Input gradients have a pivotal role in a variety of applications, including adversarial attack algorithms for evaluating model robustness, explainable AI techniques for generating saliency maps, and counterfactual explanations. However, saliency maps generated by traditional neural networks are ofte…

Cited by 10SourcePDFScholar
2023

Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex

NeurIPS 2023poster

One of the most impactful findings in computational neuroscience over the past decade is that the object recognition accuracy of deep neural networks (DNNs) correlates with their ability to predict neural responses to natural images in the inferotemporal (IT) cortex. This discovery supported the lon…

Cited by 26SourcePDFScholar
2023

Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization

NeurIPS 2023poster

Feature visualization has gained significant popularity as an explainability method, particularly after the influential work by Olah et al. in 2017. Despite its success, its widespread adoption has been limited due to issues in scaling to deeper neural networks and the reliance on tricks to generate…

Cited by 19SourcePDFScholar
2022

Harmonizing the object recognition strategies of deep neural networks with humans

NeurIPS 2022accept

The many successes of deep neural networks (DNNs) over the past decade have largely been driven by computational scale rather than insights from biological intelligence. Here, we explore if these trends have also carried concomitant improvements in explaining the visual strategies humans rely on for…

2022

Making Sense of Dependence: Efficient Black-box Explanations Using Dependence Measure

NeurIPS 2022accept

This paper presents a new efficient black-box attribution method built on Hilbert-Schmidt Independence Criterion (HSIC). Based on Reproducing Kernel Hilbert Spaces (RKHS), HSIC measures the dependence between regions of an input image and the output of a model using the kernel embedding of their dis…

2022

What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability Methods

NeurIPS 2022accept

A multitude of explainability methods has been described to try to help users better understand how modern AI systems make decisions. However, most performance metrics developed to evaluate these methods have remained largely theoretical -- without much consideration for the human end-user. In parti…

2021

Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity Analysis

NeurIPS 2021poster

We describe a novel attribution method which is grounded in Sensitivity Analysis and uses Sobol indices. Beyond modeling the individual contributions of image regions, Sobol indices provide an efficient way to capture higher-order interactions between image regions and their contributions to a neu…