← Search

Antonio D'Orazio

2 accepted papers

2026

A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Models

CVPR 2026

Despite the rapid progress in multimodal models and Large Visual-Language Models (LVLM), they remain highly susceptible to adversarial perturbations, raising serious concerns about their reliability in real-world use. While adversarial training has become the leading paradigm for building models tha

Cited by 1SourcecodeScholar
2026

Implicit Inversion turns CLIP into a Decoder

ICLR 2026poster

CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space back to images. We show that image synthesis is nevertheless p…

Cited by 0SourcecodeScholar