← Search

Mobin Bagherian

2 accepted papers

2026

Uncovering Grounding IDs: How External Cues Shape Multi-Modal Binding

ICML 2026poster

Large vision–language models (LVLMs) perform well on multimodal tasks, but their ability to reason and precisely align visual and textual information still has room for improvement. In this study, we show that external visual cues, such as symbols or grid lines, help LVLMs form more accurate connect…

Cited by 0SourceScholar
2026

Understanding Counting Mechanisms in Large Language and Vision-Language Models

CVPR 2026

Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these foundation models represent and compute numerical information in counting tasks. We use controlled experiments with repeated textual and visual items a

Cited by 0SourcecodeScholar