ACL 2025long0 citations

“What do you call a dog that is incontrovertibly true? Dogma”: Testing LLM Generalization through Humor

Alessio Cocchieri, Luca Ragazzi, Paolo Italiani, Giuseppe Tagliavini, Gianluca Moro

Abstract

Humor, requiring creativity and contextual understanding, is a hallmark of human intelligence, showcasing adaptability across linguistic scenarios. While recent advances in large language models (LLMs) demonstrate strong reasoning on various benchmarks, it remains unclear whether they truly adapt to new tasks like humans (i.e., generalize) or merely replicate memorized content. To explore this, we introduce Phunny, a new humor-based question-answering benchmark designed to assess LLMs’ reasoning through carefully crafted puns. Our dataset is manually curated to ensure novelty and minimize data contamination, providing a robust evaluation of LLMs’ linguistic comprehension. Experiments on pun comprehension, resolution, and generation reveal that most LLMs struggle with generalization, even on simple tasks, consistently underperforming the human baseline. Additionally, our detailed error analysis provides valuable insights to guide future research.

BibTeX
@inproceedings{cocchieri-etal-2025-call,
    title = "``What do you call a dog that is incontrovertibly true? Dogma'': Testing {LLM} Generalization through Humor",
    author = "Cocchieri, Alessio  and
      Ragazzi, Luca  and
      Italiani, Paolo  and
      Tagliavini, Giuseppe  and
      Moro, Gianluca",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.1117/",
    doi = "10.18653/v1/2025.acl-long.1117",
    pages = "22922--22937",
    ISBN = "979-8-89176-251-0"
}