IJCAI 20260 citations

BehaviorBench: A Psychologically Grounded Benchmark for Evaluating Personality in Large Language Models Through Realistic Behaviors

Taowen Pu, Hexi Wang, Zeyang Liu, Dongsheng Guo, Chuan Zhao

Abstract

Current approaches to evaluating personality in large language models (LLMs) typically prompt them to self-report on psychological questionnaires such as the Big Five Inventory. However, these methods assess introspective labels rather than observable behavior, despite the fact that LLMs are deployed to act in realistic contexts, not to reflect on their own traits. To bridge this gap, we introduce BehaviorBench, a new benchmark for evaluating LLM personality through concrete behaviors in everyday scenarios. Grounded in established personality psychology, BehaviorBench links each Big Five trait to validated behavioral manifestations and embeds them in contextually plausible situations that naturally elicit trait-relevant actions. We evaluate a range of models and personality shaping strategies using BehaviorBench and find a substantial mismatch between self-reported personalities and actual behaviors. By grounding evaluation in observable behavior rather than introspection, our work reveals critical gaps in current LLM personality modeling and control mechanisms. Our code and data are available at https://github.com/butra1n/BehaviorBench

Natural Language Processing: PsycholinguisticsNatural Language Processing: Resources and evaluation
BibTeX
@inproceedings{ijcai2026_behaviorbenchaps,
  title = {BehaviorBench: A Psychologically Grounded Benchmark for Evaluating Personality in Large Language Models Through Realistic Behaviors},
  author = {Taowen Pu and Hexi Wang and Zeyang Liu and Dongsheng Guo and Chuan Zhao},
  booktitle = {IJCAI 2026},
  year = {2026}
}
BehaviorBench: A Psychologically Grounded Benchmark for Evaluating Personality in Large Language Models Through Realistic Behaviors · IJCAI 2026