2025
On Guardrail Models’ Robustness to Mutations and Adversarial Attacks
EMNLP 2025
The risk of generative AI systems providing unsafe information has raised significant concerns, emphasizing the need for safety guardrails. To mitigate this risk, guardrail models are increasingly used to detect unsafe content in human-AI interactions, complementing the safety alignment of Large Lan