← Search

Alessandro Favero

9 accepted papers

2025

How Compositional Generalization and Creativity Improve as Diffusion Models are Trained

ICML 2025poster

Natural data is often organized as a hierarchical composition of features. How many samples do generative models need in order to learn the composition rules, so as to produce a combinatorially large number of novel data? What signal in the data is exploited to learn those rules? We investigate thes…

Cited by 0SourcePDFScholar
2025

LiNeS: Post-training Layer Scaling Prevents Forgetting and Enhances Model Merging

ICLR 2025poster

Fine-tuning pre-trained models has become the standard approach to endow them with specialized knowledge, but it poses fundamental challenges. In particular, (i) fine-tuning often leads to catastrophic forgetting, where improvements on a target domain degrade generalization on other tasks, and (ii)…

2025

MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs

NeurIPS 2025poster

Language models deployed in real-world systems often require post-hoc updates to incorporate new or corrected knowledge. However, editing such models efficiently and reliably—without retraining or forgetting previous information—remains a major challenge. Existing methods for lifelong model editing…

Cited by 0SourceScholar
2025

Probing the Latent Hierarchical Structure of Data via Diffusion Models

ICLR 2025poster

High-dimensional data must be highly structured to be learnable. Although the compositional and hierarchical nature of data is often put forward to explain learnability, quantitative measurements establishing these properties are scarce. Likewise, accessing the latent variables underlying such a dat…

Cited by 3SourcePDFScholar
2024

Multi-Modal Hallucination Control by Visual Information Grounding

CVPR 2024poster

Generative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers which however are not always grounded in the input image. We investigate this phenomenon usually referred to as "hallucination" and show that it stems from an excessive reliance on the language prior. I…

Cited by 72SourcePDFScholar
2023

Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models

NeurIPS 2023oral

Task arithmetic has recently emerged as a cost-effective and scalable approach to edit pre-trained models directly in weight space: By adding the fine-tuned weights of different tasks, the model's performance can be improved on these tasks, while negating them leads to task forgetting. Yet, our unde…

2023

What Can Be Learnt With Wide Convolutional Neural Networks?

ICML 2023poster

Understanding how convolutional neural networks (CNNs) can efficiently learn high-dimensional functions remains a fundamental challenge. A popular belief is that these models harness the local and hierarchical structure of natural data such as images. Yet, we lack a quantitative understanding of how…

2021

Locality defeats the curse of dimensionality in convolutional teacher-student scenarios

NeurIPS 2021poster

Convolutional neural networks perform a local and translationally-invariant treatment of the data: quantifying which of these two aspects is central to their success remains a challenge. We study this problem within a teacher-student framework for kernel regression, using 'convolutional' kernels ins…

Cited by 23SourcePDFScholar
2021

Relative stability toward diffeomorphisms indicates performance in deep nets

NeurIPS 2021poster

Understanding why deep nets can classify data in large dimensions remains a challenge. It has been proposed that they do so by becoming stable to diffeomorphisms, yet existing empirical measurements support that it is often not the case. We revisit this question by defining a maximum-entropy distrib…

Cited by 16SourcePDFScholar