2025
Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluation
EMNLP 2025
In the era of evaluating large language models (LLMs), data contamination has become an increasingly prominent concern. To address this risk, LLM benchmarking has evolved from a *static* to a *dynamic* paradigm. In this work, we conduct an in-depth analysis of existing *static* and *dynamic* benchma