2026
Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
ICML 2026poster
As frontier AI systems become increasingly capable, concerns about deceptive behaviors have intensified. Unlike hallucinations, which stem from capability limitations, deception involves strategically misleading responses despite correct internal representations. While prior work has primarily studi…