← Search

Yehan Yang

1 accepted papers

2025

From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring

NeurIPS 2025poster

Though safety alignment has been applied to most large language models (LLMs), LLM service providers generally deploy a subsequent moderation as the external safety guardrail in real-world products. Existing moderators mainly practice a conventional full detection, which determines the harmfulness b…

Cited by 0SourceScholar