EMNLP 2024finding7 citations

Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks

Jack Gallifant, Shan Chen, Pedro José Ferreira Moreira, Nikolaj Munch, Mingye Gao, Jackson Pond, Leo Anthony Celi, Hugo Aerts

Abstract

Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To study this, we create a new robustness dataset, RABBITS, to evaluate performance differences on medical benchmarks after swapping brand and generic drug names using physician expert annotations.We assess both open-source and API-based LLMs on MedQA and MedMCQA, revealing a consistent performance drop ranging from 1-10%. Furthermore, we identify a potential source of this fragility as the contamination of test data in widely used pre-training datasets.

BibTeX
@inproceedings{gallifant-etal-2024-language,
    title = "Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks",
    author = "Gallifant, Jack  and
      Chen, Shan  and
      Moreira, Pedro Jos{\'e} Ferreira  and
      Munch, Nikolaj  and
      Gao, Mingye  and
      Pond, Jackson  and
      Celi, Leo Anthony  and
      Aerts, Hugo  and
      Hartvigsen, Thomas  and
      Bitterman, Danielle",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-emnlp.726/",
    doi = "10.18653/v1/2024.findings-emnlp.726",
    pages = "12448--12465"
}
Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks · EMNLP 2024