2024
FineMatch: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
ECCV 2024poster
"Recent progress in large-scale pre-training has led to the development of advanced vision-language models (VLMs) with remarkable proficiency in comprehending and generating multimodal content. Despite the impressive ability to perform complex reasoning for VLMs, current models often struggle to eff…