← Search

Lorena Gonzalez-Manzano

2 accepted papers

2026

Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique

ICML 2026poster

We propose a system for marking sensitive or copyrighted texts to detect their use in fine-tuning large language models (LLMs) under black-box access with statistical guarantees. Our method builds digital "marks" using invisible Unicode characters organized into ("cue", "reply") pairs. During an aud…

Cited by 0SourceScholar
2025

Large language models can learn and generalize steganographic chain-of-thought under process supervision

NeurIPS 2025poster

Chain-of-thought (CoT) reasoning not only enhances large language model performance but also provides critical insights into decision-making processes, marking it as a useful tool for monitoring model intent and planning. By proactively preventing models from acting on CoT indicating misaligned or h…

Cited by 0SourceScholar