ACL 2024system demonstrations6 citations

OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety

Chuang Liu, Linhao Yu, Jiaxuan Li, Renren Jin, Yufei Huang, Ling Shi, Junhui Zhang, Xinmeng Ji

Abstract

The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluation platforms for assessing Chinese LLMs, many of these focus primarily on capabilities, usually overlooking potential alignment and safety issues. To address this gap, we introduce OpenEval, an evaluation testbed that benchmarks Chinese LLMs across capability, alignment and safety. For capability assessment, we include 12 benchmark datasets to evaluate Chinese LLMs from 4 sub-dimensions: NLP tasks, disciplinary knowledge, commonsense reasoning and mathematical reasoning. For alignment assessment, OpenEval contains 7 datasets that examines the bias, offensiveness and illegalness in the outputs yielded by Chinese LLMs. To evaluate safety, especially anticipated risks (e.g., power-seeking, self-awareness) of advanced LLMs, we include 6 datasets. In addition to these benchmarks, we have implemented a phased public evaluation and benchmark update strategy to ensure that OpenEval is in line with the development of Chinese LLMs or even able to provide cutting-edge benchmark datasets to guide the development of Chinese LLMs. In our first public evaluation, we have tested a range of Chinese LLMs, spanning from 7B to 72B parameters, including both open-source and proprietary models. Evaluation results indicate that while Chinese LLMs have shown impressive performance in certain tasks, more attention should be directed towards broader aspects such as commonsense reasoning, alignment, and safety.

BibTeX
@inproceedings{liu-etal-2024-openeval,
    title = "{O}pen{E}val: Benchmarking {C}hinese {LLM}s across Capability, Alignment and Safety",
    author = "Liu, Chuang  and
      Yu, Linhao  and
      Li, Jiaxuan  and
      Jin, Renren  and
      Huang, Yufei  and
      Shi, Ling  and
      Zhang, Junhui  and
      Ji, Xinmeng  and
      Cui, Tingting  and
      Liutao, Liutao  and
      Song, Jinwang  and
      Zan, Hongying  and
      Li, Sun  and
      Xiong, Deyi",
    editor = "Cao, Yixin  and
      Feng, Yang  and
      Xiong, Deyi",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-demos.19/",
    doi = "10.18653/v1/2024.acl-demos.19",
    pages = "190--210"
}
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety · ACL 2024