← Search

Enric Boix-Adserà

8 accepted papers

2026

FACT: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations

ICLR 2026poster

It is a central challenge in deep learning to understand how neural networks learn representations. A leading approach is the Neural Feature Ansatz (NFA) (Radhakrishnan et al., 2024), a conjectured mechanism for how feature learning occurs. Although the NFA is empirically validated, it is an educate…

Cited by 0SourceScholar
2025

Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones

NeurIPS 2025poster

Inference-time computation has emerged as a promising scaling axis for improving large language model reasoning. However, despite yielding impressive performance, the optimal allocation of inference-time computation remains poorly understood. A central question is whether to prioritize sequential sc…

Cited by 0SourcecodeScholar
2024

Prompts have evil twins

EMNLP 2024main

We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts “evil twins” because they are obfuscated and uninterpretable (evil), but at the same time mimi…

2024

When can transformers reason with abstract symbols?

ICLR 2024poster

We investigate the capabilities of transformer models on relational reasoning tasks. In these tasks, models are trained on a set of strings encoding abstract relations, and are then tested out-of-distribution on data that contains symbols that did not appear in the training dataset. We prove that fo…

2023

Transformers learn through gradual rank increase

NeurIPS 2023poster

We identify incremental learning dynamics in transformers, where the difference between trained and initial weights progressively increases in rank. We rigorously prove this occurs under the simplifying assumptions of diagonal weight matrices and small initialization. Our experiments support the the…

Cited by 26SourcePDFScholar
2022

GULP: a prediction-based metric between representations

NeurIPS 2022accept

Comparing the representations learned by different neural networks has recently emerged as a key tool to understand various architectures and ultimately optimize them. In this work, we introduce GULP, a family of distance measures between representations that is explicitly motivated by downstream p…

2021

The staircase property: How hierarchical structure can guide deep learning

NeurIPS 2021poster

This paper identifies a structural property of data distributions that enables deep neural networks to learn hierarchically. We define the ``staircase'' property for functions over the Boolean hypercube, which posits that high-order Fourier coefficients are reachable from lower-order Fourier coeffic…

Cited by 68SourcePDFScholar