← Search

Xuemeng Weng

1 accepted papers

2025

Detoxifying Large Language Models via the Diversity of Toxic Samples

EMNLP 2025

Eliminating toxicity from Large Language Models (LLMs) is crucial for ensuring user safety. However, current methods have limitations in the analysis and utilization of toxic samples, failing to fully harness their potential. Through comparative analysis of toxic and safe samples, we discover that t