2023
Evaluating GPT-3 Generated Explanations for Hateful Content Moderation
IJCAI 2023poster
Recent research has focused on using large language models (LLMs) to generate explanations for hate speech through fine-tuning or prompting. Despite the growing interest in this area, these generated explanations' effectiveness and potential limitations remain poorly understood. A key concern is tha…