← Search

Dario Loi

1 accepted papers

2026

Implicit Inversion turns CLIP into a Decoder

ICLR 2026poster

CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space back to images. We show that image synthesis is nevertheless p…

Cited by 0SourcecodeScholar