2026
Always Refuse: Steering LLMs Against Jailbreaks with Contrastive Activations (Student Abstract)
AAAI 2026technical
Refusals must be resilient, not brittle.” Yet guarding refusals against adversarial phrasing and shifting user contexts remains difficult: large language models (LLMs) still yield to jailbreak prompts that evade safety filters and surface harmful content. We propose Refusal Activation Steering (RAS)