← Search

Victor De Marez

1 accepted papers

2025

In Benchmarks We Trust ... Or Not?

EMNLP 2025

Standardized benchmarks are central to evaluating and comparing model performance in Natural Language Processing (NLP). However, Large Language Models (LLMs) have exposed shortcomings in existing benchmarks, and so far there is no clear solution. In this paper, we survey a wide scope of benchmarking

Cited by 0SourcePDFScholar