2026
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
ICML 2026oral
Large Language Models (LLMs) increasingly act as autonomous agents with tool use, ability to execute code, file I/O, and network access. These capabilities create novel security risks. To mitigate these risks, agents are often deployed and evaluated in isolated environments commonly referred to as s…