← Search

Iacopo Masi

14 accepted papers

2026

A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Models

CVPR 2026

Despite the rapid progress in multimodal models and Large Visual-Language Models (LVLM), they remain highly susceptible to adversarial perturbations, raising serious concerns about their reliability in real-world use. While adversarial training has become the leading paradigm for building models tha

Cited by 1SourcecodeScholar
2026

Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization

ICLR 2026poster

Vision–Language Models (VLMs) have become essential for tasks such as image synthesis, captioning, and retrieval by aligning textual and visual information in a shared embedding space. Yet, this flexibility also makes them vulnerable to malicious prompts designed to produce unsafe content, raising c…

Cited by 0SourceScholar
2026

Implicit Inversion turns CLIP into a Decoder

ICLR 2026poster

CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space back to images. We show that image synthesis is nevertheless p…

Cited by 0SourcecodeScholar
2026

MASS: MoErging through Adaptive Subspace Selection

ICLR 2026poster

Model merging has recently emerged as a lightweight alternative to ensembling, combining multiple fine-tuned models into a single set of parameters with no additional training overhead. Yet, existing merging methods fall short of matching the full accuracy of separately fine-tuned endpoints. We pres…

Cited by 0SourcecodeScholar
2023

Hierarchical Fine-Grained Image Forgery Detection and Localization

CVPR 2023poster

Differences in forgery attributes of images generated in CNN-synthesized and image-editing domains are large, and such differences make a unified image forgery detection and localization (IFDL) challenging. To this end, we present a hierarchical fine-grained formulation for IFDL representation learn…

2023

MAGIC: Mask-Guided Image Synthesis by Inverting a Quasi-robust Classifier

AAAI 2023technical

We offer a method for one-shot mask-guided image synthesis that allows controlling manipulations of a single image by inverting a quasi-robust classifier equipped with strong regularizers. Our proposed method, entitled MAGIC, leverages structured gradients from a pre-trained quasi-robust classifier…

2020

Towards Learning Structure via Consensus for Face Segmentation and Parsing

CVPR 2020poster

Face segmentation is the task of densely labeling pixels on the face according to their semantics. While current methods place an emphasis on developing sophisticated architectures, use conditional random fields for smoothness, or rather employ adversarial training, we follow an alternative path tow…

Cited by 21PDFcodeScholar
2020

Two-branch Recurrent Network for Isolating Deepfakes in Videos

ECCV 2020poster

The current spike of hyper-realistic faces artificially generated using deepfakes calls for media forensics solutions that are tailored to video streams and work reliably with a low false alarm rate at the video level. We present a method for deepfake detection based on a two-branch network structur…

2019

AIRD: Adversarial Learning Framework for Image Repurposing Detection

CVPR 2019poster

Image repurposing is a commonly used method for spreading misinformation on social media and online forums, which involves publishing untampered images with modified metadata to create rumors and further propaganda. While manual verification is possible, given vast amounts of verified knowledge avai…

Cited by 30PDFcodeScholar
2018

Extreme 3D Face Reconstruction: Seeing Through Occlusions

CVPR 2018poster

Existing single view, 3D face reconstruction methods can produce beautifully detailed 3D results, but typically only for near frontal, unobstructed viewpoints. We describe a system designed to provide detailed 3D reconstructions of faces viewed under extreme conditions, out of plane rotations, and o…

2017

Regressing Robust and Discriminative 3D Morphable Models With a Very Deep Neural Network

CVPR 2017poster

The 3D shapes of faces are well known to be discriminative. Yet despite this, they are rarely used for face recognition and always under controlled viewing conditions. We claim that this is a symptom of a serious but often overlooked problem with existing methods for single view 3D face reconstructi…

Cited by 612PDFScholar