← Search

Tomaso Poggio

10 accepted papers

2026

A universal compression theory: Lottery ticket hypothesis and superpolynomial scaling laws

ICLR 2026poster

When training large-scale models, the performance typically scales with the number of parameters and the dataset size according to a slow power law. A fundamental theoretical and practical question is whether comparable performance can be achieved with significantly smaller models and substantially…

Cited by 0SourceScholar
2025

Training the Untrainable: Introducing Inductive Bias via Representational Alignment

NeurIPS 2025poster

We demonstrate that architectures which traditionally are considered to be ill-suited for a task can be trained using inductive biases from another architecture. We call a network untrainable when it overfits, underfits, or converges to poor results even when tuning their hyperparameters. For examp…

Cited by 0SourceScholar
2023

Norm-based Generalization Bounds for Sparse Neural Networks

NeurIPS 2023poster

In this paper, we derive norm-based generalization bounds for sparse ReLU neural networks, including convolutional neural networks. These bounds differ from previous ones because they consider the sparse structure of the neural network architecture and the norms of the convolutional filters, rather…

Cited by 3SourcePDFScholar
2023

System Identification of Neural Systems: If We Got It Right, Would We Know?

ICML 2023poster

Artificial neural networks are being proposed as models of parts of the brain. The networks are compared to recordings of biological neurons, and good performance in reproducing neural responses is considered to support the model's validity. A key question is how much this system identification appr…

Cited by 18SourcePDFScholar
2020

Biologically Inspired Mechanisms for Adversarial Robustness

NeurIPS 2020poster

A convolutional neural network strongly robust to adversarial perturbations at reasonable computational and performance cost has not yet been demonstrated. The primate visual ventral stream seems to be robust to small perturbations in visual stimuli but the underlying mechanisms that give rise to th…

Cited by 36SourcePDFScholar
2019

Biologically-Plausible Learning Algorithms Can Scale to Large Datasets

ICLR 2019poster

The backpropagation (BP) algorithm is often thought to be biologically implausible in the brain. One of the main reasons is that BP requires symmetric weight matrices in the feedforward and feedback pathways. To address this “weight transport problem” (Grossberg, 1987), two biologically-plausible al…

2019

Fisher-Rao Metric, Geometry, and Complexity of Neural Networks

AISTATS 2019poster

We study the relationship between geometry and capacity measures for deep neural networks from an invariance viewpoint. We introduce a new notion of capacity — the Fisher-Rao norm — that possesses desirable invariance properties and is motivated by Information Geometry. We discover an analytical cha…

Cited by 274SourcePDFScholar
2015

Convex Learning of Multiple Tasks and their Structure

ICML 2015poster

Reducing the amount of human supervision is a key problem in machine learning and a natural approach is that of exploiting the relations (structure) among different tasks. This is the idea at the core of multi-task learning. In this context a fundamental question is how to incorporate the tasks stru…

Cited by 94SourcePDFScholar