← Search

Hosein Hasani

8 accepted papers

2026

HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models

ICML 2026poster

Large vision-language models can produce object hallucinations in image descriptions, highlighting the need for effective detection and mitigation strategies. Prior work commonly relies on the model's attention weights on visual tokens as a detection signal. We reveal that coarse-grained attention-b…

Cited by 0SourceScholar
2026

Uncovering Grounding IDs: How External Cues Shape Multi-Modal Binding

ICML 2026poster

Large vision–language models (LVLMs) perform well on multimodal tasks, but their ability to reason and precisely align visual and textual information still has room for improvement. In this study, we show that external visual cues, such as symbols or grid lines, help LVLMs form more accurate connect…

Cited by 0SourceScholar
2026

Understanding Counting Mechanisms in Large Language and Vision-Language Models

CVPR 2026

Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these foundation models represent and compute numerical information in counting tasks. We use controlled experiments with repeated textual and visual items a

Cited by 0SourcecodeScholar
2025

Matina: A Culturally-Aligned Persian Language Model Using Multiple LoRA Experts

ACL 2025finding

Large language models (LLMs) are powerful tools for a variety of applications, but to interact effectively with users, they must align with the cultural values and linguistic nuances of their audience. However, existing LLMs often fall short in adequately modeling underrepresented languages and cult…

Cited by 0SourcePDFScholar
2025

Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection

NeurIPS 2025poster

Out-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications, where they frequently face data distributions unseen during training. Despite progress, existing methods are often vulnerable to spurious correlations that mi…

Cited by 0SourceScholar
2025

Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs

NeurIPS 2025poster

Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This limitation underlies persistent errors in tasks such as counting, vis…

Cited by 0SourceScholar
2021

Generative vs. Discriminative: Rethinking The Meta-Continual Learning

NeurIPS 2021poster

Deep neural networks have achieved human-level capabilities in various learning tasks. However, they generally lose performance in more realistic scenarios like learning in a continual manner. In contrast, humans can incorporate their prior knowledge to learn new concepts efficiently without forgett…

2019

Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks

NeurIPS 2019poster

Numerous neurophysiological studies have revealed that a large number of the primary visual cortex neurons operate in a regime called surround modulation. Surround modulation has a substantial effect on various perceptual tasks, and it also plays a crucial role in the efficient neural coding of the…

Cited by 26SourcePDFScholar