EMNLP 2024industry0 citations

Don’t be my Doctor! Recognizing Healthcare Advice in Large Language Models

Kellen Tan Cheng, Anna Lisa Gentile, Pengyuan Li, Chad DeLuca, Guang-Jie Ren

Abstract

Large language models (LLMs) have seen increasing popularity in daily use, with their widespread adoption by many corporations as virtual assistants, chatbots, predictors, and many more. Their growing influence raises the need for safeguards and guardrails to ensure that the outputs from LLMs do not mislead or harm users. This is especially true for highly regulated domains such as healthcare, where misleading advice may influence users to unknowingly commit malpractice. Despite this vulnerability, the majority of guardrail benchmarking datasets do not focus enough on medical advice specifically. In this paper, we present the HeAL benchmark (HEalth Advice in LLMs), a health-advice benchmark dataset that has been manually curated and annotated to evaluate LLMs’ capability in recognizing health-advice - which we use to safeguard LLMs deployed in industrial settings. We use HeAL to assess several models and report a detailed analysis of the findings.

BibTeX
@inproceedings{cheng-etal-2024-dont,
    title = "Don`t be my Doctor! Recognizing Healthcare Advice in Large Language Models",
    author = "Cheng, Kellen Tan  and
      Gentile, Anna Lisa  and
      Li, Pengyuan  and
      DeLuca, Chad  and
      Ren, Guang-Jie",
    editor = "Dernoncourt, Franck  and
      Preo{\c{t}}iuc-Pietro, Daniel  and
      Shimorina, Anastasia",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track",
    month = nov,
    year = "2024",
    address = "Miami, Florida, US",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-industry.72/",
    doi = "10.18653/v1/2024.emnlp-industry.72",
    pages = "970--980"
}
Don’t be my Doctor! Recognizing Healthcare Advice in Large Language Models · EMNLP 2024