2025
Towards Robust Comparisons of NLP Models: A Case Study
COLING 2025main
Comparing the test scores of different NLP models across downstream datasets to determine which model leads to the most accurate results is the ultimate step in any experimental work. Doing so via a single mean score may not accurately quantify the real capabilities of the models. Previous works hav…