ACL 2025long0 citations

VMLU Benchmarks: A comprehensive benchmark toolkit for Vietnamese LLMs

Cuc Thi Bui, Nguyen Truong Son, Truong Van Trang, Lam Viet Phung, Pham Nhut Huy, Hoang Anh Le, Quoc Huu Van, Phong Nguyen-Thuan Do

Abstract

The evolution of Large Language Models (LLMs) has underscored the necessity for benchmarks designed for various languages and cultural contexts. To address this need for Vietnamese, we present the first Vietnamese Multitask Language Understanding (VMLU) Benchmarks. The VMLU benchmarks consist of four datasets that assess different capabilities of LLMs, including general knowledge, reading comprehension, reasoning, and conversational skills. This paper also provides an insightful overview of the current state of some dominant LLMs, such as Llama-3, Qwen2.5, and GPT-4, highlighting their performances and limitations when measured against these benchmarks. Furthermore, we provide insights into how prompt design can influence VMLU’s evaluation outcomes, as well as suggest that open-source LLMs can serve as effective, cost-efficient evaluators within the Vietnamese context. By offering a comprehensive and accessible benchmarking framework, the VMLU Benchmarks aim to foster the development and fine-tuning of Vietnamese LLMs, thereby establishing a foundation for their practical applications in language-specific domains.

BibTeX
@inproceedings{bui-etal-2025-vmlu,
    title = "{VMLU} Benchmarks: A comprehensive benchmark toolkit for {V}ietnamese {LLM}s",
    author = "Bui, Cuc Thi  and
      Son, Nguyen Truong  and
      Trang, Truong Van  and
      Phung, Lam Viet  and
      Huy, Pham Nhut  and
      Le, Hoang Anh  and
      Van, Quoc Huu  and
      Do, Phong Nguyen-Thuan  and
      Truc, Van Le Tran  and
      Chau, Duc Thanh  and
      Nguyen, Le-Minh",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.563/",
    doi = "10.18653/v1/2025.acl-long.563",
    pages = "11495--11515",
    ISBN = "979-8-89176-251-0"
}