← Search

Mir Rayat Imtiaz Hossain

4 accepted papers

2026

Segmentation From Attention: Training-Free Layer Selection and One-Shot Tuning for Segmentation in VLMs

ICML 2026poster

Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associations between textual descriptions and image regions. This emergent ability enables zero-shot object detection and segmenta…

Cited by 0SourceScholar
2025

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

EMNLP 2025

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models attend to irrelevant regions of the chart. In this work, we prese

Cited by 0SourcePDFScholar
2024

Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach

CVPR 2024poster

The emergence of attention-based transformer models has led to their extensive use in various tasks due to their superior generalization and transfer properties. Recent research has demonstrated that such models when prompted appropriately are excellent for few-shot inference. However such technique…

Cited by 9SourcePDFScholar