2025
Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation
NeurIPS 2025poster
Large Language Models have demonstrated impressive fluency across diverse tasks, yet their tendency to produce toxic content remains a critical challenge for AI safety and public trust. Existing toxicity mitigation approaches primarily manipulate individual neuron activations, but these methods suff…