← Search

Minseung Lee

3 accepted papers

2026

Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding

CVPR 2026

Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However, processing visually complex and information-rich images, such as infographics or document layouts, requires these mode

Cited by 0SourcecodeScholar