When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems
While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels. Most existing MAS security defenses are built upon two core assumptions: semantically-explicit malicious attacks …