EMNLP 2023short findings0 citations

GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captions

Ting-Yao Hsu, Chieh-Yang Huang, Ryan A. Rossi, Sungchul Kim, C. Lee Giles, Ting-Hao Kenneth Huang

Abstract

There is growing interest in systems that generate captions for scientific figures. However, assessing these systems' output poses a significant challenge. Human evaluation requires academic expertise and is costly, while automatic evaluation depends on often low-quality author-written captions. This paper investigates using large language models (LLMs) as a cost-effective, reference-free method for evaluating figure captions. We first constructed SCICAP-EVAL, a human evaluation dataset that contains human judgments for 3,600 scientific figure captions, both original and machine-made, for 600 arXiv figures. We then prompted LLMs like GPT-4 and GPT-3 to score (1-6) each caption based on its potential to aid reader understanding, given relevant context such as figure-mentioning paragraphs. Results show that GPT-4, used as a zero-shot evaluator, outperformed all other models and even surpassed assessments made by computer science undergraduates, achieving a Kendall correlation score of 0.401 with Ph.D. students' rankings.

text generationscientific figure captioncaption evaluation
BibTeX
@inproceedings{
hsu2023gpt,
title={{GPT}-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captions},
author={Ting-Yao Hsu and Chieh-Yang Huang and Ryan A. Rossi and Sungchul Kim and C. Lee Giles and Ting-Hao Kenneth Huang},
booktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},
year={2023},
url={https://openreview.net/forum?id=gVTtkPJbRq}
}
GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captions · EMNLP 2023