Break the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models
Watermarking is a principled approach for tracing the provenance of large language model (LLM) outputs, but its deployment in practice is hindered by inference inefficiency. Speculative sampling accelerates inference, with efficiency improving as the acceptance rate between draft and target models i…