2025
PolyGuard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
NeurIPS 2025poster
As large language models (LLMs) become widespread across diverse applications, concerns about the security and safety of LLM interactions have intensified. Numerous guardrail models and benchmarks have been developed to ensure LLM content safety. However, existing guardrail benchmarks are often buil…