LLM Stinger: Jailbreaking LLMs Using RL Fine-Tuned LLMs (Student Abstract)
We introduce LLM Stinger, a novel approach that leverages Large Language Models (LLMs) to automatically generate adversarial suffixes for jailbreak attacks. Unlike traditional methods, which require complex prompt engineering or white-box access, LLM Stinger uses a reinforcement learning (RL) loop t…