EMNLP 2024main0 citations

Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark

Fenglin Liu, Zheng Li, Hongjian Zhou, Qingyu Yin, Jingfeng Yang, Xianfeng Tang, Chen Luo, Ming Zeng

Abstract

The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the close-ended question-answering (QA) task with answer options for evaluation. However, many clinical decisions involve answering open-ended questions without pre-set options. To better understand LLMs in the clinic, we construct a benchmark ClinicBench. We first collect eleven existing datasets covering diverse clinical language generation, understanding, and reasoning tasks. Furthermore, we construct six novel datasets and clinical tasks that are complex but common in real-world practice, e.g., open-ended decision-making, long document processing, and emerging drug analysis. We conduct an extensive evaluation of twenty-two LLMs under both zero-shot and few-shot settings. Finally, we invite medical experts to evaluate the clinical usefulness of LLMs

BibTeX
@inproceedings{liu-etal-2024-large,
    title = "Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark",
    author = "Liu, Fenglin  and
      Li, Zheng  and
      Zhou, Hongjian  and
      Yin, Qingyu  and
      Yang, Jingfeng  and
      Tang, Xianfeng  and
      Luo, Chen  and
      Zeng, Ming  and
      Jiang, Haoming  and
      Gao, Yifan  and
      Nigam, Priyanka  and
      Nag, Sreyashi  and
      Yin, Bing  and
      Hua, Yining  and
      Zhou, Xuan  and
      Rohanian, Omid  and
      Thakur, Anshul  and
      Clifton, Lei  and
      Clifton, David A.",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.759/",
    doi = "10.18653/v1/2024.emnlp-main.759",
    pages = "13696--13710"
}
Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark · EMNLP 2024