2026
Enhancing Pre-training Data Detection in LLMs Through Discriminative and Symmetric Prefix Selection
AAAI 2026technical
The rapid development of large language models (LLMs) has relied on access to high-quality, large-scale datasets, yet growing concerns around data privacy and security have spurred substantial research into pre-training data detection. While state-of-the-art (SOTA) methods such as RECALL and CON-REC