HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims
Michiel Van Der Meer, Pavel Korshunov, Sébastien Marcel, Lonneke Van Der Plas
Abstract
Misinformation can be countered with fact-checking, but the process is costly and slow. Identifying checkworthy claims is the first step, where automation can help scale fact-checkers’ efforts. However, detection methods struggle with content that is (1) multimodal, (2) from diverse domains, and (3) synthetic. We introduce HintsOfTruth, a public dataset for multimodal checkworthiness detection with 27K real-world and synthetic image/claim pairs. The mix of real and synthetic data makes this dataset unique and ideal for benchmarking detection methods. We compare fine-tuned and prompted Large Language Models (LLMs). We find that well-configured lightweight text-based encoders perform comparably to multimodal models but the former only focus on identifying non-claim-like content. Multimodal LLMs can be more accurate but come at a significant computational cost, making them impractical for large-scale applications. When faced with synthetic data, multimodal models perform more robustly.
BibTeX
@inproceedings{van-der-meer-etal-2025-hintsoftruth,
title = "{H}ints{O}f{T}ruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims",
author = "Van Der Meer, Michiel and
Korshunov, Pavel and
Marcel, S{\'e}bastien and
Plas, Lonneke Van Der",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1510/",
doi = "10.18653/v1/2025.acl-long.1510",
pages = "31274--31291",
ISBN = "979-8-89176-251-0"
}