EMNLP 2024main10 citations

FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture

Wenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu

Abstract

Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional diversity in this domain, we introduce FoodieQA, a manually curated, fine-grained image-text dataset capturing the intricate features of food cultures across various regions in China. We evaluate vision–language Models (VLMs) and large language models (LLMs) on newly collected, unseen food images and corresponding questions. FoodieQA comprises three multiple-choice question-answering tasks where models need to answer questions based on multiple images, a single image, and text-only descriptions, respectively. While LLMs excel at text-based question answering, surpassing human accuracy, the open-sourced VLMs still fall short by 41% on multi-image and 21% on single-image VQA tasks, although closed-weights models perform closer to human levels (within 10%). Our findings highlight that understanding food and its cultural implications remains a challenging and under-explored direction.

BibTeX
@inproceedings{li-etal-2024-foodieqa,
    title = "{F}oodie{QA}: A Multimodal Dataset for Fine-Grained Understanding of {C}hinese Food Culture",
    author = "Li, Wenyan  and
      Zhang, Crystina  and
      Li, Jiaang  and
      Peng, Qiwei  and
      Tang, Raphael  and
      Zhou, Li  and
      Zhang, Weijia  and
      Hu, Guimin  and
      Yuan, Yifei  and
      S{\o}gaard, Anders  and
      Hershcovich, Daniel  and
      Elliott, Desmond",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1063/",
    doi = "10.18653/v1/2024.emnlp-main.1063",
    pages = "19077--19095"
}
FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture · EMNLP 2024