← Search

Mincheol Kwon

3 accepted papers

2026

Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding

CVPR 2026

Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However, processing visually complex and information-rich images, such as infographics or document layouts, requires these mode

Cited by 0SourcecodeScholar
2026

The Truth Stays in the Family: Enhancing Contextual Truthfulness via Inherited Heads in Model Lineages

ICML 2026poster

Recent advances in large language models (LLMs) have led to the emergence of specialized multimodal LLMs (MLLMs), forming distinct model families that share a common foundation language models. Despite this evolutionary trend, it remains unexplored whether a fundamental behavioral link exists betwee…

Cited by 0SourceScholar
2025

Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding

EMNLP 2025

Large Vision-Language Models (LVLMs) have recently shown promising results on various multimodal tasks, even achieving human-comparable performance in certain cases. Nevertheless, LVLMs remain prone to hallucinations–they often rely heavily on a single modality or memorize training data without prop

Cited by 0SourcePDFScholar