← Search

Alberto Baldrati

5 accepted papers

2026

VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions

CVPR 2026

Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates an information bottleneck that impairs performance, especially on challenging tasks that require fine-grained understan

Cited by 0SourceScholar
2025

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

ICLR 2025poster

Pre-trained multi-modal Vision-Language Models like CLIP are widely used off-the-shelf for a variety of applications. In this paper, we show that the common practice of individually exploiting the text or image encoders of these powerful multi-modal models is highly suboptimal for intra-modal tasks…

2024

Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation

ECCV 2024poster

"Vision-Language Models (VLMs) demonstrate remarkable zero-shot generalization to unseen tasks, but fall short of the performance of supervised methods in generalizing to downstream tasks with limited data. Prompt learning is emerging as a parameter-efficient method for adapting VLMs, but state-of-t…

2023

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

ICCV 2023poster

Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previo…

Cited by 75PDFcodeScholar
2023

Zero-Shot Composed Image Retrieval with Textual Inversion

ICCV 2023poster

Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption that describes the difference between the two images. The high effort and cost required for labeling datasets for CIR hamper the widespread usage of existing methods,…

Cited by 126PDFcodeScholar