2026
Causal Detection of Multi-Step LLM Agent Attacks
ICML 2026poster
Multi-step prompt injection attacks on LLM agents present a fundamental detection challenge because malicious intent emerges only after workflows complete, while individual actions remain legitimate in isolation. Existing defenses, including input sanitization, output validation, and instruction hie…