2026
Rethinking Pretraining Data Detection for LLMs: From Local to Global
ICML 2026poster
The advancements of Large Language Models (LLMs) are primarily attributed to massive pretraining data, which also introduces risks like privacy leakage and data contamination. Therefore, it is crucial to determine whether an LLM has been trained on a given target text. Existing detection methods pri…