← Search

Shunji Wan

1 accepted papers

2024

VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation

EMNLP 2024finding

As large language models achieve impressive scores on traditional benchmarks, an increasing number of researchers are becoming concerned about benchmark data leakage during pre-training, commonly known as the data contamination problem. To ensure fair evaluation, recent benchmarks release only the t…