2024
VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation
EMNLP 2024finding
As large language models achieve impressive scores on traditional benchmarks, an increasing number of researchers are becoming concerned about benchmark data leakage during pre-training, commonly known as the data contamination problem. To ensure fair evaluation, recent benchmarks release only the t…