← Search

Fatemeh Askari

3 accepted papers

2026

Uncovering Grounding IDs: How External Cues Shape Multi-Modal Binding

ICML 2026poster

Large vision–language models (LVLMs) perform well on multimodal tasks, but their ability to reason and precisely align visual and textual information still has room for improvement. In this study, we show that external visual cues, such as symbols or grid lines, help LVLMs form more accurate connect…

Cited by 0SourceScholar
2026

Understanding Counting Mechanisms in Large Language and Vision-Language Models

CVPR 2026

Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these foundation models represent and compute numerical information in counting tasks. We use controlled experiments with repeated textual and visual items a

Cited by 0SourcecodeScholar
2025

Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs

NeurIPS 2025poster

Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This limitation underlies persistent errors in tasks such as counting, vis…

Cited by 0SourceScholar