← Search

Marco Fumero

16 accepted papers

2026

Boomerang Distillation Enables Zero-Shot Model Size Interpolation

ICLR 2026poster

Large language models (LLMs) are typically deployed under diverse memory and compute constraints. Existing approaches build model families by training each size independently, which is prohibitively expensive and provides only coarse-grained size options. In this work, we identify a novel phenomenon…

Cited by 0SourcecodeScholar
2026

Identifiability and recoverability in self-supervised models

ICLR 2026poster

Self-supervised models exhibit a surprising stability in their internal representations. Whereas most prior work treats this stability as a single property, we formalize it as two distinct concepts: **statistical identifiability** (consistency of representations across runs) and **structural identi…

Cited by 0SourcecodeScholar
2026

Learning Explicit Single-Cell Dynamics Using ODE Representations

ICLR 2026poster

Modeling the dynamics of cellular differentiation is fundamental to advancing the understanding and treatment of diseases associated with this process, such as cancer. With the rapid growth of single-cell datasets, this has also become a particularly promising and active domain for machine learning.…

Cited by 0SourcecodeScholar
2025

Connecting Neural Models Latent Geometries with Relative Geodesic Representations

NeurIPS 2025poster

Neural models learn representations of high-dimensional data on low-dimensional manifolds. Multiple factors, including stochasticities in the training process, model architectures, and additional inductive biases, may induce different representations, even when learning the same task on the same dat…

Cited by 0SourcecodeScholar
2025

Unifying Causal Representation Learning with the Invariance Principle

ICLR 2025poster

Causal representation learning (CRL) aims at recovering latent causal variables from high-dimensional observations to solve causal downstream tasks, such as predicting the effect of new interventions or more robust classification. A plethora of methods have been developed, each tackling carefully…

Cited by 5SourcePDFScholar
2024

$C^2M^3$: Cycle-Consistent Multi-Model Merging

NeurIPS 2024poster

In this paper, we present a novel data-free method for merging neural networks in weight space. Our method optimizes for the permutations of network neurons while ensuring global coherence across all layers, and it outperforms recent layer-local approaches in a set of challenging scenarios. We then…

2024

From Bricks to Bridges: Product of Invariances to Enhance Latent Space Communication

ICLR 2024spotlight

It has been observed that representations learned by distinct neural networks conceal structural similarities when the models are trained under similar inductive biases. From a geometric perspective, identifying the classes of transformations and the related invariances that connect these representa…

Cited by 12SourcePDFScholar
2024

Latent Functional Maps: a spectral framework for representation alignment

NeurIPS 2024poster

Neural models learn data representations that lie on low-dimensional manifolds, yet modeling the relation between these representational spaces is an ongoing challenge. By integrating spectral geometry principles into neural modeling, we show that this problem can be better addressed in the function…

Cited by 2SourcePDFScholar
2023

ASIF: Coupled Data Turns Unimodal Models to Multimodal without Training

NeurIPS 2023poster

CLIP proved that aligning visual and language spaces is key to solving many vision tasks without explicit training, but required to train image and text encoders from scratch on a huge dataset. LiT improved this by only training the text encoder and using a pre-trained vision network. In this paper,…

Cited by 36SourcePDFScholar
2023

Latent Space Translation via Semantic Alignment

NeurIPS 2023poster

While different neural models often exhibit latent spaces that are alike when exposed to semantically related data, this intrinsic similarity is not always immediately discernible. Towards a better understanding of this phenomenon, our work shows how representations learned from these neural modules…

2023

Leveraging sparse and shared feature activations for disentangled representation learning

NeurIPS 2023spotlight

Recovering the latent factors of variation of high dimensional data has so far focused on simple synthetic settings. Mostly building on unsupervised and weakly-supervised objectives, prior work missed out on the positive implications for representation learning on real world data. In this work, we p…

Cited by 22SourcePDFScholar
2023

Relative representations enable zero-shot latent space communication

ICLR 2023top-5%

Neural networks embed the geometric structure of a data manifold lying in a high-dimensional space into latent representations. Ideally, the distribution of the data points in the latent space should depend only on the task, the data, the loss, and other architecture-specific constraints. However, f…

Cited by 101SourcePDFScholar
2022

CLIP-Forge: Towards Zero-Shot Text-To-Shape Generation

CVPR 2022poster

Generating shapes using natural language can enable new ways of imagining and creating the things around us. While significant recent progress has been made in text-to-image generation, text-to-shape generation remains a challenging problem due to the unavailability of paired text and shape data at…

Cited by 323PDFcodeScholar
2021

Learning disentangled representations via product manifold projection

ICML 2021spotlight

We propose a novel approach to disentangle the generative factors of variation underlying a given set of observations. Our method builds upon the idea that the (unknown) low-dimensional manifold underlying the data space can be explicitly modeled as a product of submanifolds. This definition of dise…

Cited by 31SourcePDFScholar