ICLR 2025poster1 citations

Chain-of-region: Visual Language Models Need Details for Diagram Analysis

Xue Li, Yiyou Sun, Wei Cheng, Yinglun Zhu, Haifeng Chen

Abstract

Visual Language Models (VLMs) like GPT-4V have broadened the scope of LLM applications, yet they face significant challenges in accurately processing visual details, particularly in scientific diagrams. This paper explores the necessity of meticulous visual detail collection and region decomposition for enhancing the performance of VLMs in scientific diagram analysis. We propose a novel approach that combines traditional computer vision techniques with VLMs to systematically decompose diagrams into discernible visual elements and aggregate essential metadata. Our method employs techniques in OpenCV library to identify and label regions, followed by a refinement process using shape detection and region merging algorithms, which are particularly suited to the structured nature of scientific diagrams. This strategy not only improves the granularity and accuracy of visual information processing but also extends the capabilities of VLMs beyond their current limitations. We validate our approach through a series of experiments that demonstrate enhanced performance in diagram analysis tasks, setting a new standard for integrating visual and language processing in a multimodal context.

Multi-modalityVisual Language ModelComputer Vision
BibTeX
@inproceedings{
li2025chainofregion,
title={Chain-of-region: Visual Language Models Need  Details for Diagram Analysis},
author={Xue Li and Yiyou Sun and Wei Cheng and Yinglun Zhu and Haifeng Chen},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=M6fYrICcQs}
}
Chain-of-region: Visual Language Models Need Details for Diagram Analysis · ICLR 2025