2025
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
NeurIPS 2025poster
LLMs have demonstrated impressive capabilities across various natural language processing tasks yet remain vulnerable to prompts, known as jailbreak attacks, carefully designed to bypass safety guardrails and elicit harmful responses. Traditional methods rely on manual heuristics that suffer from li…