2024
Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe Sampling
NeurIPS 2024poster
Safety of Large Language Models (LLMs) has become a central issue given their rapid progress and wide applications. Greedy Coordinate Gradient (GCG) is shown to be effective in constructing prompts containing adversarial suffixes to break the presumingly safe LLMs, but the optimization of GCG is tim…