← Search

Marco Bertini

7 accepted papers

2025

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

ICLR 2025poster

Pre-trained multi-modal Vision-Language Models like CLIP are widely used off-the-shelf for a variety of applications. In this paper, we show that the common practice of individually exploiting the text or image encoders of these powerful multi-modal models is highly suboptimal for intra-modal tasks…

2024

CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding

NeurIPS 2024poster

The comic domain is rapidly advancing with the development of single-page analysis and synthesis models. However, evaluation metrics and datasets lag behind, often limited to small-scale or single-style test sets. We introduce a novel benchmark, CoMix, designed to evaluate the multi-task capabilitie…

2024

Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation

ECCV 2024poster

"Vision-Language Models (VLMs) demonstrate remarkable zero-shot generalization to unseen tasks, but fall short of the performance of supervised methods in generalizing to downstream tasks with limited data. Prompt learning is emerging as a parameter-efficient method for adapting VLMs, but state-of-t…

2023

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

ICCV 2023poster

Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previo…

Cited by 75PDFcodeScholar
2023

Zero-Shot Composed Image Retrieval with Textual Inversion

ICCV 2023poster

Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption that describes the difference between the two images. The high effort and cost required for labeling datasets for CIR hamper the widespread usage of existing methods,…

Cited by 126PDFcodeScholar
2020

Task-conditioned Domain Adaptation for Pedestrian Detection in Thermal Imagery

ECCV 2020poster

Pedestrian detection is a core problem in computer vision that sees broad application in video surveillance and, more recently, in advanced driving assistance systems. Despite its broad application and interest, it remains a challenging problem in part due to the vast range of conditions under which…

2017

Deep Generative Adversarial Compression Artifact Removal

ICCV 2017poster

Compression artifacts arise in images whenever a lossy compression algorithm is applied. These artifacts eliminate details present in the original image, or add noise and small structures; because of these effects they make images less pleasant for the human eye, and may also lead to decreased perfo…

Cited by 257PDFScholar