LawBench: Benchmarking Legal Knowledge of Large Language Models
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang, Kai Chen
Abstract
We present LawBench, the first evaluation benchmark composed of 20 tasks aimed to assess the ability of Large Language Models (LLMs) to perform Chinese legal-related tasks. LawBench is meticulously crafted to enable precise assessment of LLMs’ legal capabilities from three cognitive levels that correspond to the widely accepted Bloom’s cognitive taxonomy. Using LawBench, we present a comprehensive evaluation of 21 popular LLMs and the first comparative analysis of the empirical results in order to reveal their relative strengths and weaknesses. All data, model predictions and evaluation code are accessible from https://github.com/open-compass/LawBench.
BibTeX
@inproceedings{fei-etal-2024-lawbench,
title = "{L}aw{B}ench: Benchmarking Legal Knowledge of Large Language Models",
author = "Fei, Zhiwei and
Shen, Xiaoyu and
Zhu, Dawei and
Zhou, Fengzhe and
Han, Zhuo and
Huang, Alan and
Zhang, Songyang and
Chen, Kai and
Yin, Zhixin and
Shen, Zongwen and
Ge, Jidong and
Ng, Vincent",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.emnlp-main.452/",
doi = "10.18653/v1/2024.emnlp-main.452",
pages = "7933--7962"
}