← Search

Tommaso Mencattini

3 accepted papers

2026

Language Models are Injective and Hence Invertible

ICLR 2026poster

Transformer components such as non-linear activations and normalization are inherently non-injective, suggesting that different inputs could map to the same output and prevent exact recovery of the input from a model’s representations. In this paper, we challenge this view. First, we prove mathemati…

Cited by 0SourcecodeScholar
2025

MERGE$^3$: Efficient Evolutionary Merging on Consumer-grade GPUs

ICML 2025poster

Evolutionary model merging enables the creation of high-performing multi-task models but remains computationally prohibitive for consumer hardware. We introduce MERGE$^3$, an efficient framework that makes evolutionary merging of Large Language Models (LLMs) feasible on a single GPU by reducing fitn…