2025
BAM-ICL: Causal Hijacking In-Context Learning with Budgeted Adversarial Manipulation
NeurIPS 2025poster
Recent research shows that large language models (LLMs) are vulnerable to hijacking attacks under the scenario of in-context learning (ICL) where LLMs demonstrate impressive capabilities in performing tasks by conditioning on a sequence of in-context examples (ICEs) (i.e., prompts with task-specific…