2025
Multi-Turn Jailbreaking Large Language Models via Attention Shifting
AAAI 2025technical
Large Language Models (LLMs) have achieved significant performance in various natural language processing tasks but also pose safety and ethical threats, thus requiring red teaming and alignment processes to bolster their safety. To effectively exploit these aligned LLMs, recent studies have introdu…