2025
ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
EMNLP 2025
Large Language Models (LLMs) have revolutionized Natural Language Processing (NLP) but pose risks of inadvertently exposing copyrighted or proprietary data, especially when such data is used for training but not intended for distribution. Traditional methods address these leaks only after content is