EMNLP 2024main101 citations

LawBench: Benchmarking Legal Knowledge of Large Language Models

Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang, Kai Chen

Abstract

We present LawBench, the first evaluation benchmark composed of 20 tasks aimed to assess the ability of Large Language Models (LLMs) to perform Chinese legal-related tasks. LawBench is meticulously crafted to enable precise assessment of LLMs’ legal capabilities from three cognitive levels that correspond to the widely accepted Bloom’s cognitive taxonomy. Using LawBench, we present a comprehensive evaluation of 21 popular LLMs and the first comparative analysis of the empirical results in order to reveal their relative strengths and weaknesses. All data, model predictions and evaluation code are accessible from https://github.com/open-compass/LawBench.

BibTeX
@inproceedings{fei-etal-2024-lawbench,
    title = "{L}aw{B}ench: Benchmarking Legal Knowledge of Large Language Models",
    author = "Fei, Zhiwei  and
      Shen, Xiaoyu  and
      Zhu, Dawei  and
      Zhou, Fengzhe  and
      Han, Zhuo  and
      Huang, Alan  and
      Zhang, Songyang  and
      Chen, Kai  and
      Yin, Zhixin  and
      Shen, Zongwen  and
      Ge, Jidong  and
      Ng, Vincent",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.452/",
    doi = "10.18653/v1/2024.emnlp-main.452",
    pages = "7933--7962"
}
LawBench: Benchmarking Legal Knowledge of Large Language Models · EMNLP 2024