EMNLP 2024finding4 citations

Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness

Srija Mukhopadhyay, Adnan Qidwai, Aparna Garimella, Pritika Ramu, Vivek Gupta, Dan Roth

Abstract

Chart question answering (CQA) is a crucial area of Visual Language Understanding. However, the robustness and consistency of current Visual Language Models (VLMs) in this field remain under-explored. This paper evaluates state-of-the-art VLMs on comprehensive datasets, developed specifically for this study, encompassing diverse question categories and chart formats. We investigate two key aspects: 1) the models’ ability to handle varying levels of chart and question complexity, and 2) their robustness across different visual representations of the same underlying data. Our analysis reveals significant performance variations based on question and chart types, highlighting both strengths and weaknesses of current models. Additionally, we identify areas for improvement and propose future research directions to build more robust and reliable CQA systems. This study sheds light on the limitations of current models and paves the way for future advancements in the field.

BibTeX
@inproceedings{mukhopadhyay-etal-2024-unraveling,
    title = "Unraveling the Truth: Do {VLM}s really Understand Charts? A Deep Dive into Consistency and Robustness",
    author = "Mukhopadhyay, Srija  and
      Qidwai, Adnan  and
      Garimella, Aparna  and
      Ramu, Pritika  and
      Gupta, Vivek  and
      Roth, Dan",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-emnlp.973/",
    doi = "10.18653/v1/2024.findings-emnlp.973",
    pages = "16696--16717"
}
Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness · EMNLP 2024