← Search

Hu XiaoLong

1 accepted papers

2026

Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models

ICLR 2026poster

Data contamination poses a significant threat to the reliable evaluation of Large Language Models (LLMs). This issue arises when benchmark samples may inadvertently appear in training sets, compromising the validity of reported performance. While detection methods have been developed for the pre-tra…

Cited by 0SourcecodeScholar