2025
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
ICLR 2025poster
As large language models (LLMs) grow more powerful, they also become more difficult to trust. They could be either aligned with human intentions, or exhibit "subversive misalignment" -- introducing subtle errors that bypass safety checks. Although individual errors may not immediately cause harm, ea…