2024
CLEAN–EVAL: Clean Evaluation on Contaminated Large Language Models
NAACL 2024findings
We are currently in an era of fierce competition among various large language models (LLMs), continuously pushing the boundaries of benchmark performance. However, genuinely assessing the capabilities of these LLMs has become a challenging and critical issue due to potential data contamination. In t…