← Search

Lorenzo Baraldi

27 accepted papers

2026

ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

CVPR 2026

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in jointly understanding text, images, and videos, often evaluated via Visual Question Answering (VQA). However, even state-of-the-art MLLMs struggle with domain-specific or knowledge-intensive queries, where relevant inform

Cited by 0SourcecodeScholar
2025

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

CVPR 2025poster

Multimodal LLMs (MLLMs) are the natural extension of large language models to handle multimodal inputs, combining text and image data. They have recently garnered attention due to their capability to address complex tasks involving both modalities. However, their effectiveness is limited to the know…

2025

Causal Graphical Models for Vision-Language Compositional Understanding

ICLR 2025poster

Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language, usually modeling an image caption as a “bag of words”. As a result, they perform poorly on compositional tasks, which require a deeper understanding o…

2025

Hyperbolic Safety-Aware Vision-Language Models

CVPR 2025highlight

Addressing the retrieval of unsafe content from vision-language models such as CLIP is an important step towards real-world integration. Current efforts have relied on unlearning techniques that try to erase the model's knowledge of unsafe concepts. While effective in reducing unwanted outputs, unle…

2025

MissRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models

ICCV 2025poster

Recently, Multimodal Large Language Models (MLLMs) have emerged as a leading framework for enhancing the ability of Large Language Models (LLMs) to interpret non-linguistic modalities. Despite their impressive capabilities, the robustness of MLLMs under conditions where one or more modalities are mi…

2025

Multimodal Emotion Recognition in Conversation via Possible Speaker's Audio and Visual Sequence Selection

ICASSP 2025accepted

Multimodal Emotion Recognition in Conversation (MERC) is an important element in human-machine interaction. It allows machines to automatically identify and track the emotional status of speakers during a conversation in a multimodal setting. However, the conversations involving various audio and vi…

Cited by 0SourceScholar
2025

Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval

CVPR 2025poster

Cross-modal retrieval is gaining increasing efficacy and interest from the research community, thanks to large-scale training, novel architectural and learning designs, and its application in LLMs and multimodal LLMs. In this paper, we move a step forward and design an approach that allows for multi…

2025

Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation

ICCV 2025poster

Open-Vocabulary Segmentation (OVS) aims at segmenting images from free-form textual concepts without predefined training classes. While existing vision-language models such as CLIP can generate segmentation masks by leveraging coarse spatial information from Vision Transformers, they face challenges…

2025

What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models

ICCV 2025poster

Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce…

2025

vHector and HeisenVec: Scalable Vector Graphics Generation Through Large Language Models

NeurIPS 2025poster

We introduce HeisenVec, a large-scale dataset designed to advance research in vector graphics generation from natural language descriptions. Unlike conventional image generation datasets that focus on raster images, HeisenVec targets the structured and symbolic domain of Scalable Vector Graphics (SV…

Cited by 0SourceScholar
2024

BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues

ECCV 2024poster

"Effectively aligning with human judgment when evaluating machine-generated image captions represents a complex yet intriguing challenge. Existing evaluation metrics like CIDEr or CLIP-Score fall short in this regard as they do not take into account the corresponding image or lack the capability of…

2024

Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities

ECCV 2024poster

"Discerning between authentic content and that generated by advanced AI methods has become increasingly challenging. While previous research primarily addresses the detection of fake faces, the identification of generated natural images has only recently surfaced. This prompted the recent exploratio…

2024

Mapping High-level Semantic Regions in Indoor Environments without Object Recognition

ICRA 2024poster

Robots require a semantic understanding of their surroundings to operate in an efficient and explainable way in human environments. In the literature, there has been an extensive focus on object labeling and exhaustive scene graph generation; less effort has been focused on the task of purely identi…

Cited by 5SourceScholar
2024

Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments

NeurIPS 2024poster

In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly. This growth can be attributed to the recent availability of large navigation datasets in photo-realistic simulated environments, like Gibson and Matterport3D. However, the na…

2024

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

ECCV 2024poster

"Large-scale vision-and-language models, such as CLIP, are typically trained on web-scale data, which can introduce inappropriate content and lead to the development of unsafe and biased behavior. This, in turn, hampers their applicability in sensitive and trustworthy contexts and could raise signif…

2024

The Revolution of Multimodal Large Language Models: A Survey

ACL 2024findings

Connecting text and visual modalities plays an essential role in generative intelligence. For this reason, inspired by the success of large language models, significant research efforts are being devoted to the development of Multimodal Large Language Models (MLLMs). These models can seamlessly inte…

Cited by 66SourcePDFScholar
2024

Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype Generation

CVPR 2024poster

Open-vocabulary semantic segmentation aims at segmenting arbitrary categories expressed in textual form. Previous works have trained over large amounts of image-caption pairs to enforce pixel-level multimodal alignments. However captions provide global information about the semantics of a given imag…

2023

Embodied Agents for Efficient Exploration and Smart Scene Description

ICRA 2023poster

The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step towards this objective, in this work, we tackle a setting for vi…

Cited by 8SourceScholar
2023

Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation

CVPR 2023highlight

The CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positi…

2023

With a Little Help from Your Own Past: Prototypical Memory Networks for Image Captioning

ICCV 2023poster

Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it into linguistically coherent descriptions. Although successful, the attention operator only considers a weighted summation…

Cited by 23PDFcodeScholar
2022

Focus on Impact: Indoor Exploration With Intrinsic Motivation

RA-L 2022

Exploration of indoor environments has recently experienced a significant interest, also thanks to the introduction of deep neural agents built in a hierarchical fashion and trained with Deep Reinforcement Learning (DRL) on simulated environments. Current state-of-the-art methods employ a dense extr

Cited by 23SourcecodeScholar
2020

Meshed-Memory Transformer for Image Captioning

CVPR 2020poster

Transformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored. With the aim of filling this gap, we present M2…

Cited by 1303PDFcodeScholar
2020

SMArT: Training Shallow Memory-aware Transformers for Robotic Explainability

ICRA 2020poster

The ability to generate natural language explanations conditioned on the visual perception is a crucial step towards autonomous agents which can explain themselves and communicate with humans. While the research efforts in image and video captioning are giving promising results, this is often done a…

Cited by 36SourceScholar
2019

Art2Real: Unfolding the Reality of Artworks via Semantically-Aware Image-To-Image Translation

CVPR 2019poster

The applicability of computer vision to real paintings and artworks has been rarely investigated, even though a vast heritage would greatly benefit from techniques which can understand and process data from the artistic domain. This is partially due to the small amount of annotated artistic data, wh…

Cited by 119PDFcodeScholar
2019

Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions

CVPR 2019poster

Current captioning approaches can describe images using black-box architectures whose behavior is hardly controllable and explainable from the exterior. As an image can be described in infinite ways depending on the goal and the context at hand, a higher degree of controllability is needed to apply…

Cited by 231PDFcodeScholar
2018

LAMV: Learning to Align and Match Videos With Kernelized Temporal Layers

CVPR 2018poster

This paper considers a learnable approach for comparing and aligning videos. Our architecture builds upon and revisits temporal match kernels within neural networks: we propose a new temporal layer that finds temporal alignments by maximizing the scores between two sequences of vectors, according to…

2017

Hierarchical Boundary-Aware Neural Encoder for Video Captioning

CVPR 2017poster

The use of Recurrent Neural Networks for video captioning has recently gained a lot of attention, since they can be used both to encode the input video and to generate the corresponding description. In this paper, we present a recurrent video encoding scheme which can discover and leverage the hiera…

Cited by 248PDFcodeScholar