EMNLP 2024finding0 citations

Exploring the Capability of Multimodal LLMs with Yonkoma Manga: The YManga Dataset and Its Challenging Tasks

Qi Yang, Jingjie Zeng, Liang Yang, Zhihao Yang, Hongfei Lin

Abstract

Yonkoma Manga, characterized by its four-panel structure, presents unique challenges due to its rich contextual information and strong sequential features. To address the limitations of current multimodal large language models (MLLMs) in understanding this type of data, we create a novel dataset named YManga from the Internet. After filtering out low-quality content, we collect a dataset of 1,015 yonkoma strips, containing 10,150 human annotations. We then define three challenging tasks for this dataset: panel sequence detection, generation of the author’s creative intention, and description generation for masked panels. These tasks progressively introduce the complexity of understanding and utilizing such image-text data. To the best of our knowledge, YManga is the first dataset specifically designed for yonkoma manga strips understanding. Extensive experiments conducted on this dataset reveal significant challenges faced by current multimodal large language models. Our results show a substantial performance gap between models and humans across all three tasks.

BibTeX
@inproceedings{yang-etal-2024-exploring-capability,
    title = "Exploring the Capability of Multimodal {LLM}s with Yonkoma Manga: The {YM}anga Dataset and Its Challenging Tasks",
    author = "Yang, Qi  and
      Zeng, Jingjie  and
      Yang, Liang  and
      Yang, Zhihao  and
      Lin, Hongfei",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-emnlp.506/",
    doi = "10.18653/v1/2024.findings-emnlp.506",
    pages = "8672--8687"
}
Exploring the Capability of Multimodal LLMs with Yonkoma Manga: The YManga Dataset and Its Challenging Tasks · EMNLP 2024