← Search

Alexandra Bensamoun

1 accepted papers

2026

Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique

ICML 2026poster

We propose a system for marking sensitive or copyrighted texts to detect their use in fine-tuning large language models (LLMs) under black-box access with statistical guarantees. Our method builds digital "marks" using invisible Unicode characters organized into ("cue", "reply") pairs. During an aud…

Cited by 0SourceScholar