← Search

Rita Cucchiara

56 accepted papers

2026

Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals

ICLR 2026poster

While virtual try-on (VTON) systems aim to render a garment onto a target person, this paper tackles the novel task of virtual try-off (VTOFF), which addresses the inverse problem: generating standardized product images from real-world photos of clothed individuals. Unlike VTON, which must resolve d…

Cited by 0SourcecodeScholar
2026

ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

CVPR 2026

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in jointly understanding text, images, and videos, often evaluated via Visual Question Answering (VQA). However, even state-of-the-art MLLMs struggle with domain-specific or knowledge-intensive queries, where relevant inform

Cited by 0SourcecodeScholar
2026

Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models

ICML 2026poster

Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases p…

Cited by 0SourceScholar
2026

Rethinking Expressivity and Degradation-Awareness in Attention for All-in-One Blind Image Restoration

ICLR 2026poster

All-in-one image restoration (IR) aims to recover high-quality images from diverse degradations, which in real-world settings are often mixed and unknown. Unlike single-task IR, this problem requires a model to approximate a family of heterogeneous inverse functions, making it fundamentally more cha…

Cited by 0SourceScholar
2026

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

ICML 2026poster

Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through continuous-time dynamics. However, existing flow-based editors predominantly support global or single-instruction edits and …

Cited by 0SourceScholar
2025

A Second-Order Perspective on Model Compositionality and Incremental Learning

ICLR 2025spotlight

The fine-tuning of deep pre-trained models has revealed compositional properties, with multiple specialized modules that can be arbitrarily composed into a single, multi-task model. However, identifying the conditions that promote compositionality remains an open issue, with recent efforts concentra…

2025

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

CVPR 2025poster

Multimodal LLMs (MLLMs) are the natural extension of large language models to handle multimodal inputs, combining text and image data. They have recently garnered attention due to their capability to address complex tasks involving both modalities. However, their effectiveness is limited to the know…

2025

Causal Graphical Models for Vision-Language Compositional Understanding

ICLR 2025poster

Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language, usually modeling an image caption as a “bag of words”. As a result, they perform poorly on compositional tasks, which require a deeper understanding o…

2025

Diffusion Transformers for Tabular Data Time Series Generation

ICLR 2025poster

Tabular data generation has recently attracted a growing interest due to its different application scenarios. However, generating time series of tabular data, where each element of the series depends on the others, remains a largely unexplored domain. This gap is probably due to the difficulty of…

2025

DitHub: A Modular Framework for Incremental Open-Vocabulary Object Detection

NeurIPS 2025poster

Open-Vocabulary object detectors can generalize to an unrestricted set of categories through simple textual prompting. However, adapting these models to rare classes or reinforcing their abilities on multiple specialized domains remains essential. While recent methods rely on monolithic adaptation s…

Cited by 0SourceScholar
2025

Hyperbolic Safety-Aware Vision-Language Models

CVPR 2025highlight

Addressing the retrieval of unsafe content from vision-language models such as CLIP is an important step towards real-world integration. Current efforts have relied on unlearning techniques that try to erase the model's knowledge of unsafe concepts. While effective in reducing unwanted outputs, unle…

2025

Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives

IJCAI 2025

The evaluation of machine-generated captions is a complex and evolving challenge. With the advent of Multimodal Large Language Models (MLLMs), image captioning has become a core task, increasing the need for robust and reliable evaluation metrics. This survey provides a comprehensive overview of adv

2025

MissRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models

ICCV 2025poster

Recently, Multimodal Large Language Models (MLLMs) have emerged as a leading framework for enhancing the ability of Large Language Models (LLMs) to interpret non-linguistic modalities. Despite their impressive capabilities, the robustness of MLLMs under conditions where one or more modalities are mi…

2025

Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction

ICCV 2025poster

Predicting human gaze scanpaths is crucial for understanding visual attention, with applications in human-computer interaction, autonomous systems, and cognitive robotics. While deep learning models have advanced scanpath prediction, most existing approaches generate averaged behaviors, failing to c…

2025

Multimodal Emotion Recognition in Conversation via Possible Speaker's Audio and Visual Sequence Selection

ICASSP 2025accepted

Multimodal Emotion Recognition in Conversation (MERC) is an important element in human-machine interaction. It allows machines to automatically identify and track the emotional status of speakers during a conversation in a multimodal setting. However, the conversations involving various audio and vi…

Cited by 0SourceScholar
2025

Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval

CVPR 2025poster

Cross-modal retrieval is gaining increasing efficacy and interest from the research community, thanks to large-scale training, novel architectural and learning designs, and its application in LLMs and multimodal LLMs. In this paper, we move a step forward and design an approach that allows for multi…

2025

Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation

ICCV 2025poster

Open-Vocabulary Segmentation (OVS) aims at segmenting images from free-form textual concepts without predefined training classes. While existing vision-language models such as CLIP can generate segmentation masks by leveraging coarse spatial information from Vision Transformers, they face challenges…

2025

What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models

ICCV 2025poster

Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce…

2025

Zero-Shot Styled Text Image Generation, but Make It Autoregressive

CVPR 2025poster

Styled Handwritten Text Generation (HTG) has recently received attention from the computer vision and document analysis communities, which have developed several solutions, either GAN- or diffusion-based, that achieved promising results. Nonetheless, these strategies fail to generalize to novel styl…

Cited by 0SourcePDFScholar
2024

BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues

ECCV 2024poster

"Effectively aligning with human judgment when evaluating machine-generated image captions represents a complex yet intriguing challenge. Existing evaluation metrics like CIDEr or CLIP-Score fall short in this regard as they do not take into account the corresponding image or lack the capability of…

2024

Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities

ECCV 2024poster

"Discerning between authentic content and that generated by advanced AI methods has become increasingly challenging. While previous research primarily addresses the detection of fake faces, the identification of generated natural images has only recently surfaced. This prompted the recent exploratio…

2024

Is Multiple Object Tracking a Matter of Specialization?

NeurIPS 2024poster

End-to-end transformer-based trackers have achieved remarkable performance on most human-related datasets. However, training these trackers in heterogeneous scenarios poses significant challenges, including negative interference - where the model learns conflicting scene-specific parameters - and li…

Cited by 0SourcePDFScholar
2024

Mapping High-level Semantic Regions in Indoor Environments without Object Recognition

ICRA 2024poster

Robots require a semantic understanding of their surroundings to operate in an efficient and explainable way in human environments. In the literature, there has been an extensive focus on object labeling and exhaustive scene graph generation; less effort has been focused on the task of purely identi…

Cited by 5SourceScholar
2024

Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas

ECCV 2024poster

"Diffusion models have become the State-of-the-Art for text-to-image generation, and increasing research effort has been dedicated to adapting the inference process of pretrained diffusion models to achieve zero-shot capabilities. An example is the generation of panorama images, which has been tackl…

2024

Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments

NeurIPS 2024poster

In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly. This growth can be attributed to the recent availability of large navigation datasets in photo-realistic simulated environments, like Gibson and Matterport3D. However, the na…

2024

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

ECCV 2024poster

"Large-scale vision-and-language models, such as CLIP, are typically trained on web-scale data, which can introduce inappropriate content and lead to the development of unsafe and biased behavior. This, in turn, hampers their applicability in sensitive and trustworthy contexts and could raise signif…

2024

Sharing Key Semantics in Transformer Makes Efficient Image Restoration

NeurIPS 2024poster

Image Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergence of Vision Transformers (ViTs) has further propelled these advancements. When computing, the self-attention mechanism,…

2024

The Revolution of Multimodal Large Language Models: A Survey

ACL 2024findings

Connecting text and visual modalities plays an essential role in generative intelligence. For this reason, inspired by the success of large language models, significant research efforts are being devoted to the development of Multimodal Large Language Models (MLLMs). These models can seamlessly inte…

Cited by 66SourcePDFScholar
2024

Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype Generation

CVPR 2024poster

Open-vocabulary semantic segmentation aims at segmenting arbitrary categories expressed in textual form. Previous works have trained over large amounts of image-caption pairs to enforce pixel-level multimodal alignments. However captions provide global information about the semantics of a given imag…

2024

Trends, Applications, and Challenges in Human Attention Modelling

IJCAI 2024poster

Human attention modelling has proven, in recent years, to be particularly useful not only for understanding the cognitive processes underlying visual exploration, but also for providing support to artificial intelligence models that aim to solve problems in various domains, including image and video…

2023

Embodied Agents for Efficient Exploration and Smart Scene Description

ICRA 2023poster

The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step towards this objective, in this work, we tackle a setting for vi…

Cited by 8SourceScholar
2023

Input Perturbation Reduces Exposure Bias in Diffusion Models

ICML 2023poster

Denoising Diffusion Probabilistic Models have shown an impressive generation quality although their long sampling chain leads to high computational costs. In this paper, we observe that a long sampling chain also leads to an error accumulation phenomenon, which is similar to the exposure bias proble…

2023

Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision Transformers

CVPR 2023poster

Position Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This cave…

2023

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

ICCV 2023poster

Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previo…

Cited by 75PDFcodeScholar
2023

Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation

CVPR 2023highlight

The CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positi…

2023

TrackFlow: Multi-Object tracking with Normalizing Flows

ICCV 2023poster

The field of multi-object tracking has recently seen a renewed interest in the good old schema of tracking-by-detection, as its simplicity and strong priors spare it from the complex design and painful babysitting of tracking-by-attention approaches. In view of this, we aim at extending tracking-by-…

Cited by 16PDFScholar
2023

With a Little Help from Your Own Past: Prototypical Memory Networks for Image Captioning

ICCV 2023poster

Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it into linguistically coherent descriptions. Although successful, the attention operator only considers a weighted summation…

Cited by 23PDFcodeScholar
2022

Dress Code: High-Resolution Multi-Category Virtual Try-On

ECCV 2022poster

"Image-based virtual try-on strives to transfer the appearance of a clothing item onto the image of a target person. Prior work focuses mainly on upper-body clothes (e.g. t-shirts, shirts, and tops) and neglects full-body or lower-body items. This shortcoming arises from a main factor: current publi…

2022

Focus on Impact: Indoor Exploration With Intrinsic Motivation

RA-L 2022

Exploration of indoor environments has recently experienced a significant interest, also thanks to the introduction of deep neural agents built in a hierarchical fashion and trained with Deep Reinforcement Learning (DRL) on simulated environments. Current state-of-the-art methods employ a dense extr

Cited by 23SourcecodeScholar
2022

How Many Observations Are Enough? Knowledge Distillation for Trajectory Forecasting

CVPR 2022poster

Accurate prediction of future human positions is an essential task for modern video-surveillance systems. Current state-of-the-art models usually rely on a "history" of past tracked locations (e.g., 3 to 5 seconds) to predict a plausible sequence of future locations (e.g., up to the next 5 seconds).…

Cited by 78PDFScholar
2022

Maximum Class Separation as Inductive Bias in One Matrix

NeurIPS 2022accept

Maximizing the separation between classes constitutes a well-known inductive bias in machine learning and a pillar of many traditional algorithms. By default, deep networks are not equipped with this inductive bias and therefore many alternative solutions have been proposed through differential opti…

2021

MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?

ICCV 2021poster

Deep learning-based methods for video pedestrian detection and tracking require large volumes of training data to achieve good performance. However, data acquisition in crowded public environments raises data privacy concerns -- we are not allowed to simply record and store data without the explicit…

Cited by 159PDFScholar
2020

Compressed Volumetric Heatmaps for Multi-Person 3D Pose Estimation

CVPR 2020poster

In this paper we present a novel approach for bottom-up multi-person 3D human pose estimation from monocular RGB images. We propose to use high resolution volumetric heatmaps to model joint locations, devising a simple and effective compression method to drastically reduce the size of this represent…

Cited by 123PDFcodeScholar
2020

Conditional Channel Gated Networks for Task-Aware Continual Learning

CVPR 2020oral

Convolutional Neural Networks experience catastrophic forgetting when optimized on a sequence of learning problems: as they meet the objective of the current training examples, their performance on previous tasks drops drastically. In this work, we introduce a novel framework to tackle this problem…

Cited by 270PDFScholar
2020

Meshed-Memory Transformer for Image Captioning

CVPR 2020poster

Transformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored. With the aim of filling this gap, we present M2…

Cited by 1303PDFcodeScholar
2020

SMArT: Training Shallow Memory-aware Transformers for Robotic Explainability

ICRA 2020poster

The ability to generate natural language explanations conditioned on the visual perception is a crucial step towards autonomous agents which can explain themselves and communicate with humans. While the research efforts in image and video captioning are giving promising results, this is often done a…

Cited by 36SourceScholar
2019

Art2Real: Unfolding the Reality of Artworks via Semantically-Aware Image-To-Image Translation

CVPR 2019poster

The applicability of computer vision to real paintings and artworks has been rarely investigated, even though a vast heritage would greatly benefit from techniques which can understand and process data from the artistic domain. This is partially due to the small amount of annotated artistic data, wh…

Cited by 119PDFcodeScholar
2019

Classifying Signals on Irregular Domains via Convolutional Cluster Pooling

AISTATS 2019poster

We present a novel and hierarchical approach for supervised classification of signals spanning over a fixed graph, reflecting shared properties of the dataset. To this end, we introduce a Convolutional Cluster Pooling layer exploiting a multi-scale clustering in order to highlight, at different reso…

Cited by 12SourcePDFScholar
2019

Latent Space Autoregression for Novelty Detection

CVPR 2019poster

Novelty detection is commonly referred as the discrimination of observations that do not conform to a learned model of regularity. Despite its importance in different application settings, designing a novelty detector is utterly complex due to the unpredictable nature of novelties and its inaccessib…

Cited by 619PDFcodeScholar
2019

Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions

CVPR 2019poster

Current captioning approaches can describe images using black-box architectures whose behavior is hardly controllable and explainable from the exterior. As an image can be described in infinite ways depending on the goal and the context at hand, a higher degree of controllability is needed to apply…

Cited by 231PDFcodeScholar
2018

LAMV: Learning to Align and Match Videos With Kernelized Temporal Layers

CVPR 2018poster

This paper considers a learnable approach for comparing and aligning videos. Our architecture builds upon and revisits temporal match kernels within neural networks: we propose a new temporal layer that finds temporal alignments by maximizing the scores between two sequences of vectors, according to…

2018

Learning to Detect and Track Visible and Occluded Body Joints in a Virtual World

ECCV 2018poster

Multi-People Tracking in an open-world setting requires a special effort in precise detection. Moreover, temporal continuity in the detection phase gains more importance when scene cluttering introduces the challenging problems of occluded targets. For the purpose, we propose a deep network architec…

Cited by 226SourcePDFScholar
2017

Hierarchical Boundary-Aware Neural Encoder for Video Captioning

CVPR 2017poster

The use of Recurrent Neural Networks for video captioning has recently gained a lot of attention, since they can be used both to encode the input video and to generate the corresponding description. In this paper, we present a recurrent video encoding scheme which can discover and leverage the hiera…

Cited by 248PDFcodeScholar