← Search

Marcella Cornia

26 accepted papers

2026

Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals

ICLR 2026poster

While virtual try-on (VTON) systems aim to render a garment onto a target person, this paper tackles the novel task of virtual try-off (VTOFF), which addresses the inverse problem: generating standardized product images from real-world photos of clothed individuals. Unlike VTON, which must resolve d…

Cited by 0SourcecodeScholar
2026

ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

CVPR 2026

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in jointly understanding text, images, and videos, often evaluated via Visual Question Answering (VQA). However, even state-of-the-art MLLMs struggle with domain-specific or knowledge-intensive queries, where relevant inform

Cited by 0SourcecodeScholar
2025

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

CVPR 2025poster

Multimodal LLMs (MLLMs) are the natural extension of large language models to handle multimodal inputs, combining text and image data. They have recently garnered attention due to their capability to address complex tasks involving both modalities. However, their effectiveness is limited to the know…

2025

Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives

IJCAI 2025

The evaluation of machine-generated captions is a complex and evolving challenge. With the advent of Multimodal Large Language Models (MLLMs), image captioning has become a core task, increasing the need for robust and reliable evaluation metrics. This survey provides a comprehensive overview of adv

2025

MissRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models

ICCV 2025poster

Recently, Multimodal Large Language Models (MLLMs) have emerged as a leading framework for enhancing the ability of Large Language Models (LLMs) to interpret non-linguistic modalities. Despite their impressive capabilities, the robustness of MLLMs under conditions where one or more modalities are mi…

2025

Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction

ICCV 2025poster

Predicting human gaze scanpaths is crucial for understanding visual attention, with applications in human-computer interaction, autonomous systems, and cognitive robotics. While deep learning models have advanced scanpath prediction, most existing approaches generate averaged behaviors, failing to c…

2025

Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval

CVPR 2025poster

Cross-modal retrieval is gaining increasing efficacy and interest from the research community, thanks to large-scale training, novel architectural and learning designs, and its application in LLMs and multimodal LLMs. In this paper, we move a step forward and design an approach that allows for multi…

2025

Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation

ICCV 2025poster

Open-Vocabulary Segmentation (OVS) aims at segmenting images from free-form textual concepts without predefined training classes. While existing vision-language models such as CLIP can generate segmentation masks by leveraging coarse spatial information from Vision Transformers, they face challenges…

2025

What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models

ICCV 2025poster

Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce…

2024

BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues

ECCV 2024poster

"Effectively aligning with human judgment when evaluating machine-generated image captions represents a complex yet intriguing challenge. Existing evaluation metrics like CIDEr or CLIP-Score fall short in this regard as they do not take into account the corresponding image or lack the capability of…

2024

Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities

ECCV 2024poster

"Discerning between authentic content and that generated by advanced AI methods has become increasingly challenging. While previous research primarily addresses the detection of fake faces, the identification of generated natural images has only recently surfaced. This prompted the recent exploratio…

2024

Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments

NeurIPS 2024poster

In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly. This growth can be attributed to the recent availability of large navigation datasets in photo-realistic simulated environments, like Gibson and Matterport3D. However, the na…

2024

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

ECCV 2024poster

"Large-scale vision-and-language models, such as CLIP, are typically trained on web-scale data, which can introduce inappropriate content and lead to the development of unsafe and biased behavior. This, in turn, hampers their applicability in sensitive and trustworthy contexts and could raise signif…

2024

The Revolution of Multimodal Large Language Models: A Survey

ACL 2024findings

Connecting text and visual modalities plays an essential role in generative intelligence. For this reason, inspired by the success of large language models, significant research efforts are being devoted to the development of Multimodal Large Language Models (MLLMs). These models can seamlessly inte…

Cited by 66SourcePDFScholar
2024

Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype Generation

CVPR 2024poster

Open-vocabulary semantic segmentation aims at segmenting arbitrary categories expressed in textual form. Previous works have trained over large amounts of image-caption pairs to enforce pixel-level multimodal alignments. However captions provide global information about the semantics of a given imag…

2024

Trends, Applications, and Challenges in Human Attention Modelling

IJCAI 2024poster

Human attention modelling has proven, in recent years, to be particularly useful not only for understanding the cognitive processes underlying visual exploration, but also for providing support to artificial intelligence models that aim to solve problems in various domains, including image and video…

2023

Embodied Agents for Efficient Exploration and Smart Scene Description

ICRA 2023poster

The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step towards this objective, in this work, we tackle a setting for vi…

Cited by 8SourceScholar
2023

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

ICCV 2023poster

Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previo…

Cited by 75PDFcodeScholar
2023

Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation

CVPR 2023highlight

The CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positi…

2023

With a Little Help from Your Own Past: Prototypical Memory Networks for Image Captioning

ICCV 2023poster

Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it into linguistically coherent descriptions. Although successful, the attention operator only considers a weighted summation…

Cited by 23PDFcodeScholar
2022

Dress Code: High-Resolution Multi-Category Virtual Try-On

ECCV 2022poster

"Image-based virtual try-on strives to transfer the appearance of a clothing item onto the image of a target person. Prior work focuses mainly on upper-body clothes (e.g. t-shirts, shirts, and tops) and neglects full-body or lower-body items. This shortcoming arises from a main factor: current publi…

2022

Focus on Impact: Indoor Exploration With Intrinsic Motivation

RA-L 2022

Exploration of indoor environments has recently experienced a significant interest, also thanks to the introduction of deep neural agents built in a hierarchical fashion and trained with Deep Reinforcement Learning (DRL) on simulated environments. Current state-of-the-art methods employ a dense extr

Cited by 23SourcecodeScholar
2020

Meshed-Memory Transformer for Image Captioning

CVPR 2020poster

Transformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored. With the aim of filling this gap, we present M2…

Cited by 1303PDFcodeScholar
2020

SMArT: Training Shallow Memory-aware Transformers for Robotic Explainability

ICRA 2020poster

The ability to generate natural language explanations conditioned on the visual perception is a crucial step towards autonomous agents which can explain themselves and communicate with humans. While the research efforts in image and video captioning are giving promising results, this is often done a…

Cited by 36SourceScholar
2019

Art2Real: Unfolding the Reality of Artworks via Semantically-Aware Image-To-Image Translation

CVPR 2019poster

The applicability of computer vision to real paintings and artworks has been rarely investigated, even though a vast heritage would greatly benefit from techniques which can understand and process data from the artistic domain. This is partially due to the small amount of annotated artistic data, wh…

Cited by 119PDFcodeScholar
2019

Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions

CVPR 2019poster

Current captioning approaches can describe images using black-box architectures whose behavior is hardly controllable and explainable from the exterior. As an image can be described in infinite ways depending on the goal and the context at hand, a higher degree of controllability is needed to apply…

Cited by 231PDFcodeScholar