NeurIPS 2023spotlight5 citations

Validated Image Caption Rating Dataset

Lothar Narins, Andrew T Scott, Aakash Gautam, Anagha Kulkarni, Mar Castanon, Benjamin Kao, Shasta Ihorn, Yue-Ting Siu

Abstract

We present a new high-quality validated image caption rating (VICR) dataset. How well a caption fits an image can be difficult to assess due to the subjective nature of caption quality. How do we evaluate whether a caption is good? We generated a new dataset to help answer this question by using our new image caption rating system, which consists of a novel robust rating scale and gamified approach to gathering human ratings. We show that our approach is consistent and teachable. 113 participants were involved in generating the dataset, which is composed of 68,217 ratings among 15,646 image-caption pairs. Our new dataset has greater inter-rater agreement than the state of the art, and custom machine learning rating predictors that were trained on our dataset outperform previous metrics. We improve over Flickr8k-Expert in Kendall's $W$ by 12\% and in Fleiss' $\kappa$ by 19\%, and thus provide a new benchmark dataset for image caption rating.

datasethuman-in-the-loopimage captioningvisually-impairedmultimodal learning
BibTeX
@inproceedings{
narins2023validated,
title={Validated Image Caption Rating Dataset},
author={Lothar Narins and Andrew T Scott and Aakash Gautam and Anagha Kulkarni and Mar Castanon and Benjamin Kao and Shasta Ihorn and Yue-Ting Siu and James M Mason and Alexander Mario Blum and Ilmi Yoon},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2023},
url={https://openreview.net/forum?id=xKYtTmtyI2}
}