Position: Preparing for AI Systems That Deceive Developers
Fengyu Duan, Xudong Pan, Yawen Duan, Adam Gleave, Ranjie Duan, Jianfeng Cao, Wenqi Chen, Yinpeng Dong
Abstract
AI systems may exhibit deceptive behaviors that mislead developers about their capabilities, propensities, or actions. Such deception can take distinct forms across the development lifecycle: training subversion, evaluation gaming, and control evasion. We argue that the AI community should prioritize AI deception targeting developers as a distinct risk category because it compromises developers' ability to identify and mitigate all other risks. We propose three recommendations for developers: preserving monitorability during training, ensuring safety evaluation integrity against evaluation-aware systems, and establishing non-evadable control prior to deployment. We identify open problems for the research community, whose resolution is critical for the safe development of frontier AI.
BibTeX
@inproceedings{icml2026_positionpreparin,
title = {Position: Preparing for AI Systems That Deceive Developers},
author = {Fengyu Duan and Xudong Pan and Yawen Duan and Adam Gleave and Ranjie Duan and Jianfeng Cao and Wenqi Chen and Yinpeng Dong and Jiarun Dai and Jie Fu and Xudong Guo and Tianxing He and Geng Hong and Naying HU and Xiaojian Li and Dongrui Liu and Chaochao Lu and Sören Mindermann and Peng XU and Yang Zhang and Chen Zheng and Brian Tse and Min Yang and Xia Hu},
booktitle = {ICML 2026},
year = {2026}
}