2026
Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives
ICLR 2026poster
As AI systems become more capable of complex agentic tasks, they also become more capable of pursuing undesirable objectives and causing harm. Previous work has attempted to catch these unsafe instances by interrogating models directly about their objectives and behaviors. However, the main weakness…