← Search

Hexi Wang

2 accepted papers

2026

BehaviorBench: A Psychologically Grounded Benchmark for Evaluating Personality in Large Language Models Through Realistic Behaviors

IJCAI 2026

Current approaches to evaluating personality in large language models (LLMs) typically prompt them to self-report on psychological questionnaires such as the Big Five Inventory. However, these methods assess introspective labels rather than observable behavior, despite the fact that LLMs are deploye

Cited by 0Scholar
2026

Investigating Prosocial Behavior Theory in LLM Agents Under Policy-Induced Inequities

AAAI 2026technical

As large language models (LLMs) increasingly operate as autonomous agents in social contexts, evaluating their capacity for prosocial behavior is both theoretically and practically critical. However, existing research has primarily relied on static, economically framed paradigms, lacking models that

Cited by 0SourcePDFScholar