← Search

Abdullah Mazhar

1 accepted papers

2025

Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation

NeurIPS 2025poster

Large Language Models have demonstrated impressive fluency across diverse tasks, yet their tendency to produce toxic content remains a critical challenge for AI safety and public trust. Existing toxicity mitigation approaches primarily manipulate individual neuron activations, but these methods suff…

Cited by 0SourceScholar