2026
BehaviorBench: A Psychologically Grounded Benchmark for Evaluating Personality in Large Language Models Through Realistic Behaviors
IJCAI 2026
Current approaches to evaluating personality in large language models (LLMs) typically prompt them to self-report on psychological questionnaires such as the Big Five Inventory. However, these methods assess introspective labels rather than observable behavior, despite the fact that LLMs are deploye