2025
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
NeurIPS 2025poster
Though safety alignment has been applied to most large language models (LLMs), LLM service providers generally deploy a subsequent moderation as the external safety guardrail in real-world products. Existing moderators mainly practice a conventional full detection, which determines the harmfulness b…