← Search

Mohammadali Banayeeanzade

2 accepted papers

2025

CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation

CVPR 2025poster

Contrastive Language-Image Pre-training (CLIP) models excel in zero-shot classification, yet face challenges in complex multi-object scenarios. This study offers a comprehensive analysis of CLIP's limitations in these contexts using a specialized dataset, ComCO, designed to evaluate CLIP's encoders…

2025

Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs

NeurIPS 2025poster

Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This limitation underlies persistent errors in tasks such as counting, vis…

Cited by 0SourceScholar