DETCP: Self-Detoxifying Language Models With Contrastive Pairs
Influenced by context such as tone, emotion and demographic, pre-trained language models may generate harmful text, which limits their widespread application. While detoxifying language models seeks to reduce the likelihood of generating harmful content. There are two categories of detoxification st…