← Search

Emanuele Rodolà

34 accepted papers

2026

Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success

ICML 2026poster

Model merging combines knowledge from separately fine-tuned models, yet success factors remain poorly understood. While recent work treats mergeability as an intrinsic property, we show with an architecture-agnostic framework that it fundamentally depends on both the merging method and the partner t…

Cited by 0SourceScholar
2026

EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding

ICASSP 2026poster

Audio codecs power discrete music generative modelling, music streaming and immersive media by shrinking PCM audio to bandwidth-friendly bit-rates. Recent works have gravitated towards processing in the spectral domain; however, spectrogram-domains typically struggle with phase modeling which is nat…

Cited by 0SourcePDFScholar
2026

Implicit Inversion turns CLIP into a Decoder

ICLR 2026poster

CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space back to images. We show that image synthesis is nevertheless p…

Cited by 0SourcecodeScholar
2026

Language Models are Injective and Hence Invertible

ICLR 2026poster

Transformer components such as non-linear activations and normalization are inherently non-injective, suggesting that different inputs could map to the same output and prevent exact recovery of the input from a model’s representations. In this paper, we challenge this view. First, we prove mathemati…

Cited by 0SourcecodeScholar
2026

MASS: MoErging through Adaptive Subspace Selection

ICLR 2026poster

Model merging has recently emerged as a lightweight alternative to ensembling, combining multiple fine-tuned models into a single set of parameters with no additional training overhead. Yet, existing merging methods fall short of matching the full accuracy of separately fine-tuned endpoints. We pres…

Cited by 0SourcecodeScholar
2026

Multi-Way Representation Alignment

ICML 2026poster

The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces. However, current strategies for mapping these representations are inherently pairwise, scaling quadratically with the number of models and failing to yield a con…

Cited by 0SourceScholar
2026

PHALAR: Phasors for Learned Musical Audio Representations

ICML 2026poster

Stem retrieval, the task of matching missing stems to a given audio submix, is a key challenge currently limited by models that discard temporal information. We introduce PHALAR, a contrastive framework achieving a relative accuracy increase of up to $\sim 70\%$ over the state-of-the-art while requi…

Cited by 0SourceScholar
2026

Video Unlearning via Low-Rank Refusal Vector

ICLR 2026poster

Video generative models achieve high-quality synthesis from natural-language prompts by leveraging large-scale web data. However, this training paradigm inherently exposes them to unsafe biases and harmful concepts, introducing the risk of generating undesirable or illicit content. To mitigate unsaf…

Cited by 0SourcecodeScholar
2025

COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations

ICASSP 2025accepted

We present COCOLA (Coherence-Oriented Contrastive Learning for Audio), a contrastive learning method for musical audio representations that captures the harmonic and rhythmic coherence between samples. Our method operates at the level of the individual stems composing music tracks and can input feat…

Cited by 0SourceScholar
2025

Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces

CVPR 2025poster

Recent works have shown that, when trained at scale, uni-modal 2D vision and text encoders converge to learned features that share remarkable structural properties, despite arising from different representations. However, the role of 3D encoders with respect to other modalities remains unexplored. F…

Cited by 0SourcePDFScholar
2025

Implicit-ARAP: Efficient Handle-Guided Neural Field Deformation via Local Patch Meshing

NeurIPS 2025poster

Neural fields have emerged as a powerful representation for 3D geometry, enabling compact and continuous modeling of complex shapes. Despite their expressive power, manipulating neural fields in a controlled and accurate manner -- particularly under spatial constraints -- remains an open challenge,…

Cited by 0SourceScholar
2025

MERGE$^3$: Efficient Evolutionary Merging on Consumer-grade GPUs

ICML 2025poster

Evolutionary model merging enables the creation of high-performing multi-task models but remains computationally prohibitive for consumer hardware. We introduce MERGE$^3$, an efficient framework that makes evolutionary merging of Large Language Models (LLMs) feasible on a single GPU by reducing fitn…

2025

Naturalistic Music Decoding from EEG Data via Latent Diffusion Models

ICASSP 2025accepted

In this article, we explore the potential of using latent diffusion models, a family of powerful generative models, for the task of reconstructing naturalistic music from electroencephalogram (EEG) recordings. Unlike simpler music with limited timbres, such as MIDI-generated tunes or monophonic piec…

Cited by 0SourceScholar
2025

Task Singular Vectors: Reducing Task Interference in Model Merging

CVPR 2025poster

Task Arithmetic has emerged as a simple yet effective method to merge models without additional training. However, by treating entire networks as flat parameter vectors, it overlooks key structural information and is susceptible to task interference. In this paper, we study task vectors at the layer…

2025

Update Your Transformer to the Latest Release: Re-Basin of Task Vectors

ICML 2025poster

Foundation models serve as the backbone for numerous specialized models developed through fine-tuning. However, when the underlying pretrained model is updated or retrained (e.g., on larger and more curated datasets), the fine-tuned model becomes obsolete, losing its utility and requiring retraining…

2024

$C^2M^3$: Cycle-Consistent Multi-Model Merging

NeurIPS 2024poster

In this paper, we present a novel data-free method for merging neural networks in weight space. Our method optimizes for the permutations of network neurons while ensuring global coherence across all layers, and it outperforms recent layer-local approaches in a set of challenging scenarios. We then…

2024

From Bricks to Bridges: Product of Invariances to Enhance Latent Space Communication

ICLR 2024spotlight

It has been observed that representations learned by distinct neural networks conceal structural similarities when the models are trained under similar inductive biases. From a geometric perspective, identifying the classes of transformations and the related invariances that connect these representa…

Cited by 12SourcePDFScholar
2024

Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models

ICASSP 2024accepted

Multi-Source Diffusion Models (MSDM) allow for compositional musical generation tasks: generating a set of coherent sources, creating accompaniments, and performing source separation. Despite their versatility, they require estimating the joint distribution over the sources, necessitating pre-separa…

Cited by 0SourceScholar
2024

Latent Functional Maps: a spectral framework for representation alignment

NeurIPS 2024poster

Neural models learn data representations that lie on low-dimensional manifolds, yet modeling the relation between these representational spaces is an ongoing challenge. By integrating spectral geometry principles into neural modeling, we show that this problem can be better addressed in the function…

Cited by 2SourcePDFScholar
2024

Multi-Source Diffusion Models for Simultaneous Music Generation and Separation

ICLR 2024oral

In this work, we define a diffusion-based generative model capable of both music generation and source separation by learning the score of the joint probability density of sources sharing a context. Alongside the classic total inference tasks (i.e., generating a mixture, separating the sources), we…

2024

Syncfusion: Multimodal Onset-Synchronized Video-to-Audio Foley Synthesis

ICASSP 2024accepted

Sound design involves creatively selecting, recording, and editing sound effects for various media like cinema, video games, and virtual/augmented reality. One of the most time-consuming steps when designing sound is synchronizing audio with video. In some cases, environmental recordings from video…

Cited by 0SourceScholar
2023

ASIF: Coupled Data Turns Unimodal Models to Multimodal without Training

NeurIPS 2023poster

CLIP proved that aligning visual and language spaces is key to solving many vision tasks without explicit training, but required to train image and text encoders from scratch on a huge dataset. LiT improved this by only training the text encoder and using a pre-trained vision network. In this paper,…

Cited by 36SourcePDFScholar
2023

Latent Autoregressive Source Separation

AAAI 2023technical

Autoregressive models have achieved impressive results over a wide range of domains in terms of generation quality and downstream task performance. In the continuous domain, a key factor behind this success is the usage of quantized latent spaces (e.g., obtained via VQ-VAE autoencoders), which allow…

2023

Latent Space Translation via Semantic Alignment

NeurIPS 2023poster

While different neural models often exhibit latent spaces that are alike when exposed to semantically related data, this intrinsic similarity is not always immediately discernible. Towards a better understanding of this phenomenon, our work shows how representations learned from these neural modules…

2023

Leveraging sparse and shared feature activations for disentangled representation learning

NeurIPS 2023spotlight

Recovering the latent factors of variation of high dimensional data has so far focused on simple synthetic settings. Mostly building on unsupervised and weakly-supervised objectives, prior work missed out on the positive implications for representation learning on real world data. In this work, we p…

Cited by 22SourcePDFScholar
2023

Relative representations enable zero-shot latent space communication

ICLR 2023top-5%

Neural networks embed the geometric structure of a data manifold lying in a high-dimensional space into latent representations. Ideally, the distribution of the data points in the latent space should depend only on the task, the data, the loss, and other architecture-specific constraints. However, f…

Cited by 101SourcePDFScholar
2022

3D Human Pose Estimation Using Möbius Graph Convolutional Networks

ECCV 2022poster

"3D human pose estimation is fundamental to understanding human behavior. Recently, promising results have been achieved by graph convolutional networks(GCNs), which achieve state-of-the-art performance and provide rather light-weight architectures. However, a major limitation of GCNs is their inabi…

Cited by 29SourcePDFScholar
2022

Reduced Representation of Deformation Fields for Effective Non-rigid Shape Matching

NeurIPS 2022accept

In this work we present a novel approach for computing correspondences between non-rigid objects, by exploiting a reduced representation of deformation fields. Different from existing works that represent deformation fields by training a general-purpose neural network, we advocate for an approximati…

2021

Shape Registration in the Time of Transformers

NeurIPS 2021poster

In this paper, we propose a transformer-based procedure for the efficient registration of non-rigid 3D point clouds. The proposed approach is data-driven and adopts for the first time the transformers architecture in the registration task. Our method is general and applies to different settings. Gi…

2020

LIMP: Learning Latent Shape Representations with Metric Preservation Priors

ECCV 2020poster

In this paper, we advocate the adoption of metric preservation as a powerful prior for learning latent representations of deformable 3D shapes. Key to our construction is the introduction of a geometric distortion criterion, defined directly on the decoded shapes, translating the preservation of the…

Cited by 87SourcePDFScholar
2020

Towards Precise Completion of Deformable Shapes

ECCV 2020poster

According to Aristotle, a philosopher in Ancient Greece, {\it ``the whole is greater than the sum of its parts''}. This statement was adopted to explain human perception by the Gestalt psychology school of thought in the twentieth century. Here, we claim that observing part of an object which was pr…

2016

Learning shape correspondence with anisotropic convolutional neural networks

NeurIPS 2016poster

Convolutional neural networks have achieved extraordinary results in many computer vision and pattern recognition applications; however, their adoption in the computer graphics and geometry processing communities is limited due to the non-Euclidean structure of their data. In this paper, we propose…

Cited by 638SourcePDFScholar