Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
The generation of toxic content by large language models (LLMs) remains a critical challenge for the safe deployment of language technology. We propose a novel framework for implicit knowledge editing and controlled text generation by fine-tuning LLMs with a prototype-based contrastive perplexity ob…