2026
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
ICLR 2026poster
Current safety evaluations of language models rely on benchmark-based assessments that may miss targeted vulnerabilities. We present RepIt, a simple and data-efficient framework for isolating concept-specific representations in LM activations. While existing steering methods already achieve high att…