← Search

Qisheng Su

1 accepted papers

2025

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage

EMNLP 2025

Large Language Models (LLMs) have revolutionized Natural Language Processing (NLP) but pose risks of inadvertently exposing copyrighted or proprietary data, especially when such data is used for training but not intended for distribution. Traditional methods address these leaks only after content is