NeurIPS 2025poster0 citations

AI Testing Should Account for Sophisticated Strategic Behaviour

Vojtech Kovarik, Eric Olav Chen, Sami Petersen, Alexis Ghersengorin, Vincent Conitzer

Abstract

This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility that AI systems understand their circumstances and reason strategically. Second, game-theoretic analysis can inform evaluation design by formalising and scrutinising the reasoning in evaluation-based safety cases. Drawing on examples from existing AI systems, a review of relevant research, and formal strategic analysis of a stylised evaluation scenario, we present evidence for these claims and motivate several research directions.

AI evaluationAI testingschemingdeceptionevaluation awareness
BibTeX
@inproceedings{
kovarik2025ai,
title={{AI} Testing Should Account for Sophisticated Strategic Behaviour},
author={Vojtech Kovarik and Eric Olav Chen and Sami Petersen and Alexis Ghersengorin and Vincent Conitzer},
booktitle={The Thirty-Ninth Annual Conference on Neural Information Processing Systems Position Paper Track},
year={2025},
url={https://openreview.net/forum?id=OMc0BYxND4}
}