2025
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
EMNLP 2025
Intent detection, a core component of natural language understanding, has considerably evolved as a crucial mechanism in safeguarding large language models (LLMs). While prior work has applied intent detection to enhance LLMs’ moderation guardrails, showing a significant success against content-leve