← Search

Tomaso A Poggio

10 accepted papers

2026

Momentum Further Constrains Sharpness at the Edge of Stochastic Stability

ICML 2026poster

Recent work suggests that (stochastic) gradient descent self-organizes near the instability boundary, shaping both optimization and the solutions found. Momentum and mini-batch gradients are widely used in practical deep learning optimization, but it remains unclear whether they operate in a compara…

Cited by 0SourceScholar
2026

Unraveling Syntax: Language Modeling and the Substructure of Grammars

ICML 2026spotlight

While large models achieve impressive results, their learning dynamics are far from understood. Many domains of interest -- such as natural language syntax, coding languages, arithmetic problems -- are captured by context-free grammars (CFGs). In this work, we extend prior work on neural language mo…

Cited by 0SourceScholar
2025

Position: A Theory of Deep Learning Must Include Compositional Sparsity

ICML 2025poster

Overparametrized Deep Neural Networks (DNNs) have demonstrated remarkable success in a wide variety of domains too high-dimensional for classical shallow networks subject to the curse of dimensionality. However, open questions about fundamental principles, that govern the learning dynamics of DNNs,…

Cited by 0SourcePDFScholar
2024

On the Power of Decision Trees in Auto-Regressive Language Modeling

NeurIPS 2024poster

Originally proposed for handling time series data, Auto-regressive Decision Trees (ARDTs) have not yet been explored for language modeling. This paper delves into both the theoretical and practical applications of ARDTs in this new context. We theoretically demonstrate that ARDTs can compute complex…

Cited by 0SourcePDFScholar
2023

Feature learning in deep classifiers through Intermediate Neural Collapse

ICML 2023poster

In this paper, we conduct an empirical study of the feature learning process in deep classifiers. Recent research has identified a training phenomenon called Neural Collapse (NC), in which the top-layer feature embeddings of samples from the same class tend to concentrate around their means, and the…

Cited by 51SourcePDFScholar
2015

Learning with a Wasserstein Loss

NeurIPS 2015poster

Learning to predict multi-label outputs is challenging, but in many problems there is a natural metric on the outputs that can be used to improve predictions. In this paper we develop a loss function for multi-label learning, based on the Wasserstein distance. The Wasserstein distance provides a nat…

Cited by 773SourcePDFScholar