2025
LionGuard: A Contextualized Moderation Classifier to Tackle Localized Unsafe Content
COLING 2025industry
As large language models (LLMs) become increasingly prevalent in a wide variety of applications, concerns about the safety of their outputs have become more significant. Most efforts at safety-tuning or moderation today take on a predominantly Western-centric view of safety, especially for toxic, ha…