2026
PRISON: Unmasking the Criminal Potential of Large Language Models
ICLR 2026poster
As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research has overlooked the systematic assessment of LLMs’ criminal potential in realistic interactions, where criminal potential is defined as the risk of producing harmful behavi…