← Search

Santiago Arias

1 accepted papers

2025

Gandalf the Red: Adaptive Security for LLMs

ICML 2025poster

Current evaluations of defenses against prompt attacks in large language model (LLM) applications often overlook two critical factors: the dynamic nature of adversarial behavior and the usability penalties imposed on legitimate users by restrictive defenses. We propose D-SEC (Dynamic Security Utilit…