← Search

Thomas Serre

35 accepted papers

2026

Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility

ICLR 2026poster

Language models (LMs) are used for a diverse range of tasks, from question answering to writing fantastical stories. In order to reliably accomplish these tasks, LMs must be able to discern the modal category of a sentence (i.e., whether it describes something that is possible, impossible, completel…

Cited by 0SourceScholar
2025

Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement

ICASSP 2025accepted

Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of…

Cited by 0SourceScholar
2025

Follow the Energy, Find the Path: Riemannian Metrics from Energy-Based Models

NeurIPS 2025poster

What is the shortest path between two data points lying in a high-dimensional space? While the answer is trivial in Euclidean geometry, it becomes significantly more complex when the data lies on a curved manifold—requiring a Riemannian metric to describe the space's local curvature. Estimating such…

Cited by 0SourceScholar
2025

The 3D-PC: a benchmark for visual perspective taking in humans and machines

ICLR 2025poster

Visual perspective taking (VPT) is the ability to perceive and reason about the perspectives of others. It is an essential feature of human intelligence, which develops over the first decade of life and requires an ability to process the 3D structure of visual scenes. A growing number of reports hav…

2025

Tracking objects that change in appearance with phase synchrony

ICLR 2025poster

Objects we encounter often change appearance as we interact with them. Changes in illumination (shadows), object pose, or the movement of non-rigid objects can drastically alter available image features. How do biological visual systems track objects as they change? One plausible mechanism involves…

Cited by 1SourcePDFScholar
2024

Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects

NeurIPS 2024poster

Though vision transformers (ViTs) have achieved state-of-the-art performance in a variety of settings, they exhibit surprising failures when performing tasks involving visual relations. This begs the question: how do ViTs attempt to perform tasks that require computing visual relations between objec…

2024

Latent Representation Matters: Human-like Sketches in One-shot Drawing Tasks

NeurIPS 2024poster

Humans can effortlessly draw new categories from a single exemplar, a feat that has long posed a challenge for generative models. However, this gap has started to close with recent advances in diffusion models. This one-shot drawing task requires powerful inductive biases that have not been systemat…

Cited by 0SourcePDFScholar
2024

RTify: Aligning Deep Neural Networks with Human Behavioral Decisions

NeurIPS 2024poster

Current neural network models of primate vision focus on replicating overall levels of behavioral accuracy, often neglecting perceptual decisions' rich, dynamic nature. Here, we introduce a novel computational framework to model the dynamics of human behavioral choices by learning to align the tempo…

2024

Saliency strikes back: How filtering out high frequencies improves white-box explanations

ICML 2024poster

Attribution methods correspond to a class of explainability methods (XAI) that aim to assess how individual inputs contribute to a model's decision-making process. We have identified a significant limitation in one type of attribution methods, known as ``white-box" methods. Although highly efficient…

Cited by 4SourcePDFScholar
2024

Understanding Visual Feature Reliance through the Lens of Complexity

NeurIPS 2024poster

Recent studies suggest that deep learning models' inductive bias towards favoring simpler features may be an origin of shortcut learning. Yet, there has been limited focus on understanding the complexities of the myriad features that models learn. In this work, we introduce a new metric for quantify…

Cited by 5SourcePDFScholar
2023

A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation

NeurIPS 2023spotlight

In recent years, concept-based approaches have emerged as some of the most promising explainability methods to help us interpret the decisions of Artificial Neural Networks (ANNs). These methods seek to discover intelligible visual ``concepts'' buried within the complex patterns of ANN activations i…

Cited by 56SourcePDFScholar
2023

Break It Down: Evidence for Structural Compositionality in Neural Networks

NeurIPS 2023spotlight

Though modern neural networks have achieved impressive performance in both vision and language tasks, we know little about the functions that they implement. One possibility is that neural networks implicitly break down complex tasks into subroutines, implement modular solutions to these subroutines…

2023

CRAFT: Concept Recursive Activation FacTorization for Explainability

CVPR 2023poster

Attribution methods are a popular class of explainability methods that use heatmaps to depict the most important areas of an image that drive a model decision. Nevertheless, recent work has shown that these methods have limited utility in practice, presumably because they only highlight the most sal…

2023

Computing a human-like reaction time metric from stable recurrent vision models

NeurIPS 2023spotlight

The meteoric rise in the adoption of deep neural networks as computational models of vision has inspired efforts to ``align” these models with humans. One dimension of interest for alignment includes behavioral choices, but moving beyond characterizing choice patterns to capturing temporal aspects o…

2023

Diffusion Models as Artists: Are we Closing the Gap between Humans and Machines?

ICML 2023oral

An important milestone for AI is the development of algorithms that can produce drawings that are indistinguishable from those of humans. Here, we adapt the ''diversity vs. recognizability'' scoring framework from Boutin et al (2022) and find that one-shot diffusion models have indeed started to clo…

2023

Don't Lie to Me! Robust and Efficient Explainability With Verified Perturbation Analysis

CVPR 2023poster

A variety of methods have been proposed to try to explain how deep neural networks make their decisions. Key to those approaches is the need to sample the pixel space efficiently in order to derive importance maps. However, it has been shown that the sampling methods used to date introduce biases an…

Cited by 42SourcePDFScholar
2023

Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex

NeurIPS 2023poster

One of the most impactful findings in computational neuroscience over the past decade is that the object recognition accuracy of deep neural networks (DNNs) correlates with their ability to predict neural responses to natural images in the inferotemporal (IT) cortex. This discovery supported the lon…

Cited by 26SourcePDFScholar
2023

Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization

NeurIPS 2023poster

Feature visualization has gained significant popularity as an explainability method, particularly after the influential work by Olah et al. in 2017. Despite its success, its widespread adoption has been limited due to issues in scaling to deeper neural networks and the reliance on tricks to generate…

Cited by 19SourcePDFScholar
2022

A Benchmark for Compositional Visual Reasoning

NeurIPS 2022accept

A fundamental component of human vision is our ability to parse complex visual scenes and judge the relations between their constituent objects. AI benchmarks for visual reasoning have driven rapid progress in recent years with state-of-the-art systems now reaching human accuracy on some of these be…

2022

Diversity vs. Recognizability: Human-like generalization in one-shot generative models

NeurIPS 2022accept

Robust generalization to new concepts has long remained a distinctive feature of human intelligence. However, recent progress in deep generative models has now led to neural architectures capable of synthesizing novel instances of unknown visual concepts from a single training example. Yet, a more p…

2022

Harmonizing the object recognition strategies of deep neural networks with humans

NeurIPS 2022accept

The many successes of deep neural networks (DNNs) over the past decade have largely been driven by computational scale rather than insights from biological intelligence. Here, we explore if these trends have also carried concomitant improvements in explaining the visual strategies humans rely on for…

2022

What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability Methods

NeurIPS 2022accept

A multitude of explainability methods has been described to try to help users better understand how modern AI systems make decisions. However, most performance metrics developed to evaluate these methods have remained largely theoretical -- without much consideration for the human end-user. In parti…

2021

Go with the flow: Adaptive control for Neural ODEs

ICLR 2021poster

Despite their elegant formulation and lightweight memory cost, neural ordinary differential equations (NODEs) suffer from known representational limitations. In particular, the single flow learned by NODEs cannot express all homeomorphisms from a given data space to itself, and their static weight p…

Cited by 15SourcePDFScholar
2021

Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity Analysis

NeurIPS 2021poster

We describe a novel attribution method which is grounded in Sensitivity Analysis and uses Sobol indices. Beyond modeling the individual contributions of image regions, Sobol indices provide an efficient way to capture higher-order interactions between image regions and their contributions to a neu…

2021

Tracking Without Re-recognition in Humans and Machines

NeurIPS 2021poster

Imagine trying to track one particular fruitfly in a swarm of hundreds. Higher biological visual systems have evolved to track moving objects by relying on both their appearance and their motion trajectories. We investigate if state-of-the-art spatiotemporal deep neural networks are capable of the s…

Cited by 16SourcePDFScholar
2020

Stable and expressive recurrent vision models

NeurIPS 2020spotlight

Primate vision depends on recurrent processing for reliable perception. A growing body of literature also suggests that recurrent connections improve the learning efficiency and generalization of vision models on classic computer vision challenges. Why then, are current large-scale challenges domina…

2018

Learning long-range spatial dependencies with horizontal gated recurrent units

NeurIPS 2018poster

Progress in deep learning has spawned great successes in many engineering applications. As a prime example, convolutional neural networks, a type of feedforward neural networks, are now approaching -- and sometimes even surpassing -- human accuracy on a variety of visual recognition tasks. Here, how…

2016

How Deep is the Feature Analysis underlying Rapid Visual Categorization?

NeurIPS 2016poster

Rapid categorization paradigms have a long history in experimental psychology: Characterized by short presentation times and speeded behavioral responses, these tasks highlight the efficiency with which our visual system processes natural object categories. Previous studies have shown that feed-forw…