← Search

Francesco Cagnetta

8 accepted papers

2026

Deep networks learn to parse uniform-depth context-free languages from local statistics

ICML 2026poster

Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal representations of Large Language Models (LLMs) support their ability to parse text when predicting the next word, while repre…

Cited by 0SourceScholar
2026

Deriving Neural Scaling Laws from the Statistics of Natural Language

ICML 2026poster

Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first su…

Cited by 0SourceScholar
2025

How Compositional Generalization and Creativity Improve as Diffusion Models are Trained

ICML 2025poster

Natural data is often organized as a hierarchical composition of features. How many samples do generative models need in order to learn the composition rules, so as to produce a combinatorially large number of novel data? What signal in the data is exploited to learn those rules? We investigate thes…

Cited by 0SourcePDFScholar
2025

Learning curves theory for hierarchically compositional data with power-law distributed features

ICML 2025poster

Recent theories suggest that Neural Scaling Laws arise whenever the task is linearly decomposed into units that are power-law distributed. Alternatively, scaling laws also emerge when data exhibit a hierarchically compositional structure, as is thought to occur in language and images. To unify these…

Cited by 1SourcePDFScholar
2024

Towards a theory of how the structure of language is acquired by deep neural networks

NeurIPS 2024poster

How much data is required to learn the structure of a language via next-token prediction? We study this question for synthetic datasets generated via a Probabilistic Context-Free Grammar (PCFG)---a hierarchical generative model that captures the tree-like structure of natural languages. We determine…

Cited by 12SourcePDFScholar
2023

What Can Be Learnt With Wide Convolutional Neural Networks?

ICML 2023poster

Understanding how convolutional neural networks (CNNs) can efficiently learn high-dimensional functions remains a fundamental challenge. A popular belief is that these models harness the local and hierarchical structure of natural data such as images. Yet, we lack a quantitative understanding of how…

2022

Learning sparse features can lead to overfitting in neural networks

NeurIPS 2022accept

It is widely believed that the success of deep networks lies in their ability to learn a meaningful representation of the features of the data. Yet, understanding when and how this feature learning improves performance remains a challenge: for example, it is beneficial for modern architectures train…

2021

Locality defeats the curse of dimensionality in convolutional teacher-student scenarios

NeurIPS 2021poster

Convolutional neural networks perform a local and translationally-invariant treatment of the data: quantifying which of these two aspects is central to their success remains a challenge. We study this problem within a teacher-student framework for kernel regression, using 'convolutional' kernels ins…

Cited by 23SourcePDFScholar