From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification Framework
The impressive performance of large language models (LLMs) also brings inherent toxicity risks, prompting the need for effective detoxification to support responsible deployment. Prevailing methods generally follow an inflexible model-specific fashion, addressing only individual models or model fami