← Search

Federico Cocchi

5 accepted papers

2026

ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

CVPR 2026

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in jointly understanding text, images, and videos, often evaluated via Visual Question Answering (VQA). However, even state-of-the-art MLLMs struggle with domain-specific or knowledge-intensive queries, where relevant inform

Cited by 0SourcecodeScholar
2025

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

CVPR 2025poster

Multimodal LLMs (MLLMs) are the natural extension of large language models to handle multimodal inputs, combining text and image data. They have recently garnered attention due to their capability to address complex tasks involving both modalities. However, their effectiveness is limited to the know…

2024

Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities

ECCV 2024poster

"Discerning between authentic content and that generated by advanced AI methods has become increasingly challenging. While previous research primarily addresses the detection of fake faces, the identification of generated natural images has only recently surfaced. This prompted the recent exploratio…

2024

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

ECCV 2024poster

"Large-scale vision-and-language models, such as CLIP, are typically trained on web-scale data, which can introduce inappropriate content and lead to the development of unsafe and biased behavior. This, in turn, hampers their applicability in sensitive and trustworthy contexts and could raise signif…

2024

The Revolution of Multimodal Large Language Models: A Survey

ACL 2024findings

Connecting text and visual modalities plays an essential role in generative intelligence. For this reason, inspired by the success of large language models, significant research efforts are being devoted to the development of Multimodal Large Language Models (MLLMs). These models can seamlessly inte…

Cited by 66SourcePDFScholar