ACL 2021short104 citations

X-Fact: A New Benchmark Dataset for Multilingual Fact Checking

Ashim Gupta, Vivek Srikumar

Abstract

In this work, we introduce : the largest publicly available multilingual dataset for factual verification of naturally existing real-world claims. The dataset contains short statements in 25 languages and is labeled for veracity by expert fact-checkers. The dataset includes a multilingual evaluation benchmark that measures both out-of-domain generalization, and zero-shot capabilities of the multilingual models. Using state-of-the-art multilingual transformer-based models, we develop several automated fact-checking models that, along with textual claims, make use of additional metadata and evidence from news stories retrieved using a search engine. Empirically, our best model attains an F-score of around 40%, suggesting that our dataset is a challenging benchmark for the evaluation of multilingual fact-checking models.

BibTeX
@inproceedings{gupta-srikumar-2021-x,
    title = "{X}-Fact: A New Benchmark Dataset for Multilingual Fact Checking",
    author = "Gupta, Ashim  and
      Srikumar, Vivek",
    editor = "Zong, Chengqing  and
      Xia, Fei  and
      Li, Wenjie  and
      Navigli, Roberto",
    booktitle = "Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)",
    month = aug,
    year = "2021",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.acl-short.86/",
    doi = "10.18653/v1/2021.acl-short.86",
    pages = "675--682"
}