← Search

Yuke Mei

2 accepted papers

2025

DCR: Quantifying Data Contamination in LLMs Evaluation

EMNLP 2025

The rapid advancement of large language models (LLMs) has heightened concerns about benchmark data contamination (BDC), where models inadvertently memorize evaluation data during the training process, inflating performance metrics, and undermining genuine generalization assessment. This paper introd

2025

SSA: Semantic Contamination of LLM-Driven Fake News Detection

EMNLP 2025

Benchmark data contamination (BDC) silently inflate the evaluation performance of large language models (LLMs), yet current work on BDC has centered on direct token overlap (data/label level), leaving the subtler and equally harmful semantic level BDC largely unexplored. This gap is critical in fake