2024
To the Cutoff... and Beyond? A Longitudinal Perspective on LLM Data Contamination
ICLR 2024poster
Recent claims about the impressive abilities of large language models (LLMs) are often supported by evaluating publicly available benchmarks. Since LLMs train on wide swaths of the internet, this practice raises concerns of data contamination, i.e., evaluating on examples that are explicitly or imp…