2025
Model-Dependent Moderation: Inconsistencies in Hate Speech Detection Across LLM-based Systems
ACL 2025finding
Content moderation systems powered by large language models (LLMs) are increasingly deployed to detect hate speech; however, no systematic comparison exists between different systems. If different systems produce different outcomes for the same content, it undermines consistency and predictability,…