ACL 2025finding0 citations

Can Vision Language Models Understand Mimed Actions?

Hyundong Justin Cho, Spencer Lin, Tejas Srinivasan, Michael Saxon, Deuksin Kwon, Natali T. Chavez, Jonathan May

Abstract

Non-verbal communication (NVC) is an integral part of human language, but it has been overlooked in natural language processing research. Studying NVC in general is challenging because of its high variance in interpretation among individuals and cultures, but mime—the theatrical technique of suggesting intent using only gesture, expression, and movement—is a subset of NVC with much lower human interpretation variance. As a gateway for evaluating vision-language models on their understanding of NVC, we propose Mime Identification-based Multimodal Evaluation (MIME), a gesture recognition task built upon a novel corpus of mimed activity comprising 86 unique gestures with a variety of perturbations applied to the avatar, background, and viewpoint for evaluating recognition robustness. We find that both open-weight and API-based vision-language models perform significantly worse than humans at identifying mimed gestures in MIME, motivating the need for increased research for instilling more robust understanding of human actions for VLMs.

BibTeX
@inproceedings{cho-etal-2025-vision,
    title = "Can Vision Language Models Understand Mimed Actions?",
    author = "Cho, Hyundong Justin  and
      Lin, Spencer  and
      Srinivasan, Tejas  and
      Saxon, Michael  and
      Kwon, Deuksin  and
      Chavez, Natali T.  and
      May, Jonathan",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.1372/",
    doi = "10.18653/v1/2025.findings-acl.1372",
    pages = "26744--26759",
    ISBN = "979-8-89176-256-5"
}
Can Vision Language Models Understand Mimed Actions? · ACL 2025