2026
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
ICML 2026poster
Deploying guardrails for custom policies remains challenging, as generic safety models fail to capture task-specific requirements, while prompting LLMs suffers from inconsistent boundary-case performance and high inference costs. Training custom classifiers achieves both accuracy and efficiency, yet…