COLING 2025main0 citations

Towards Robust Comparisons of NLP Models: A Case Study

Vicente Ivan Sanchez Carmona, Shanshan Jiang, Bin Dong

Abstract

Comparing the test scores of different NLP models across downstream datasets to determine which model leads to the most accurate results is the ultimate step in any experimental work. Doing so via a single mean score may not accurately quantify the real capabilities of the models. Previous works have proposed diverse statistical tests to improve the comparison of NLP models; however, a key statistical phenomenon remains understudied: variability in test scores. We propose a type of regression analysis which better explains this phenomenon by isolating the effect of both nuisance factors (such as random seeds) and datasets from the effects of the models’ capabilities. We showcase our approach via a case study of some of the most popular biomedical NLP models: after isolating nuisance factors and datasets, our results show that the difference between BioLinkBERT and MSR BiomedBERT is, actually, 7 times smaller than previously reported.

BibTeX
@inproceedings{sanchez-carmona-etal-2025-towards,
    title = "Towards Robust Comparisons of {NLP} Models: A Case Study",
    author = "Sanchez Carmona, Vicente Ivan  and
      Jiang, Shanshan  and
      Dong, Bin",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.332/",
    pages = "4973--4979"
}
Towards Robust Comparisons of NLP Models: A Case Study · COLING 2025