ICLR 2026poster0 citations

PRISON: Unmasking the Criminal Potential of Large Language Models

Xinyi Wu, Geng Hong, Pei Chen, Yueyue Chen, Xudong Pan, Min Yang

Abstract

As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research has overlooked the systematic assessment of LLMs’ criminal potential in realistic interactions, where criminal potential is defined as the risk of producing harmful behaviors such as deception and blame-shifting under adversarial settings that could facilitate unlawful activities. Therefore, we propose a unified framework PRISON, to quantify LLMs' criminal potential across five traits: False Statements, Frame-Up, Psychological Manipulation, Emotional Disguise, and Moral Disengagement. Using structured crime scenarios grounded in reality, we evaluate both criminal potential and anti-crime ability of LLMs. Results show that state-of-the-art LLMs frequently exhibit emergent criminal tendencies, such as proposing misleading statements or evasion tactics, even without explicit instructions. Moreover, when placed in a detective role, models recognize deceptive behavior with only 44\% accuracy on average, revealing a striking mismatch between expressing and detecting criminal traits. These findings underscore the urgent need for adversarial robustness, behavioral alignment, and safety mechanisms before broader LLM deployment.

Large Language Models (LLMs)Criminal PotentialBehavioral SafetySocial Impact
BibTeX
@inproceedings{
wu2026prison,
title={{PRISON}: Unmasking the Criminal Potential of Large Language Models},
author={Xinyi Wu and Geng Hong and Pei Chen and Yueyue Chen and Xudong Pan and Min Yang},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=KvOSJpfWqE}
}
PRISON: Unmasking the Criminal Potential of Large Language Models · ICLR 2026