← Search

Shir Rozenfeld

1 accepted papers

2026

GAVEL: Towards Rule-Based Safety through Activation Monitoring

ICLR 2026poster

Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. However, existing activation safety approaches, trained on broad misuse datasets, struggle with poor precision, limited fl…

Cited by 0SourcecodeScholar