2025
Exploiting Prompt-induced Confidence for Black-Box Attacks on LLMs
EMNLP 2025
Large language models (LLMs) are vulnerable to adversarial attacks even in strict black-box settings with only hard-label feedback.Existing attacks suffer from inefficient search due to lack of informative signals such as logits or probabilities. In this work, we propose Prompt-Guided Ensemble Attac