2025
Decoding Hate: Exploring Language Models’ Reactions to Hate Speech
NAACL 2025long
Hate speech is a harmful form of online expression, often manifesting as derogatory posts. It is a significant risk in digital environments. With the rise of Large Language Models (LLMs), there is concern about their potential to replicate hate speech patterns, given their training on vast amounts o…