COLING 2020main61 citations

BioMedBERT: A Pre-trained Biomedical Language Model for QA and IR

Souradip Chakraborty, Ekaba Bisong, Shweta Bhatt, Thomas Wagner, Riley Elliott, Francesco Mosconi

Abstract

The SARS-CoV-2 (COVID-19) pandemic spotlighted the importance of moving quickly with biomedical research. However, as the number of biomedical research papers continue to increase, the task of finding relevant articles to answer pressing questions has become significant. In this work, we propose a textual data mining tool that supports literature search to accelerate the work of researchers in the biomedical domain. We achieve this by building a neural-based deep contextual understanding model for Question-Answering (QA) and Information Retrieval (IR) tasks. We also leverage the new BREATHE dataset which is one of the largest available datasets of biomedical research literature, containing abstracts and full-text articles from ten different biomedical literature sources on which we pre-train our BioMedBERT model. Our work achieves state-of-the-art results on the QA fine-tuning task on BioASQ 5b, 6b and 7b datasets. In addition, we observe superior relevant results when BioMedBERT embeddings are used with Elasticsearch for the Information Retrieval task on the intelligently formulated BioASQ dataset. We believe our diverse dataset and our unique model architecture are what led us to achieve the state-of-the-art results for QA and IR tasks.

BibTeX
@inproceedings{chakraborty-etal-2020-biomedbert,
    title = "{B}io{M}ed{BERT}: A Pre-trained Biomedical Language Model for {QA} and {IR}",
    author = "Chakraborty, Souradip  and
      Bisong, Ekaba  and
      Bhatt, Shweta  and
      Wagner, Thomas  and
      Elliott, Riley  and
      Mosconi, Francesco",
    editor = "Scott, Donia  and
      Bel, Nuria  and
      Zong, Chengqing",
    booktitle = "Proceedings of the 28th International Conference on Computational Linguistics",
    month = dec,
    year = "2020",
    address = "Barcelona, Spain (Online)",
    publisher = "International Committee on Computational Linguistics",
    url = "https://aclanthology.org/2020.coling-main.59/",
    doi = "10.18653/v1/2020.coling-main.59",
    pages = "669--679"
}
BioMedBERT: A Pre-trained Biomedical Language Model for QA and IR · COLING 2020