Filtering Resistant Large Language Model Watermarking via Style Injection
The exorbitant cost of training Large Language Models (LLMs) makes it essential to protect the models from illegal copying and unauthorized usage. Recent attempts at LLM protection utilize black-box watermarking schemes, which embed distinctive input-output mapping (i.e., trigger set) directly into…