COLING 2025main0 citations

Empirical Study on Data Attributes Insufficiency of Evaluation Benchmarks for LLMs

Chuang Liu, Renren Jin, Zheng Yao, Tianyi Li, Liang Cheng, Mark Steedman, Deyi Xiong

Abstract

Previous benchmarks for evaluating large language models (LLMs) have primarily emphasized quantitative metrics, such as data volume. However, this focus may neglect key qualitative data attributes that can significantly impact the final rankings of LLMs, resulting in unreliable leaderboards. In this paper, we investigate whether current LLM benchmarks adequately consider these data attributes. We specifically examine three attributes: diversity, redundancy, and difficulty. To explore these attributes, we propose a framework with three separate modules, each designed to assess one of the attributes. Using a method that progressively incorporates these attributes, we analyze their influence on the benchmark. Our experimental results reveal a meaningful correlation between LLM rankings on the revised benchmark and the original benchmark when these attributes are accounted for. These findings indicate that existing benchmarks often fail to meet all three criteria, highlighting a lack of consideration for multifaceted data attributes in current evaluation datasets.

BibTeX
@inproceedings{liu-etal-2025-empirical,
    title = "Empirical Study on Data Attributes Insufficiency of Evaluation Benchmarks for {LLM}s",
    author = "Liu, Chuang  and
      Jin, Renren  and
      Yao, Zheng  and
      Li, Tianyi  and
      Cheng, Liang  and
      Steedman, Mark  and
      Xiong, Deyi",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.403/",
    pages = "6024--6038"
}
Empirical Study on Data Attributes Insufficiency of Evaluation Benchmarks for LLMs · COLING 2025