2026
*MemPot*: Defend Against Memory Extraction Attack with Optimized Honeypots
ICML 2026poster
Large Language Model (LLM)-based agents employ external and internal memory systems to handle complex, goal-oriented tasks, yet this exposes them to severe extraction attacks, and corresponding defenses are currently lacking. In this paper, we propose *MemPot*, the first theoretically verified defen…