COLING 2025main2 citations

Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits

Bohan Li, Jiannan Guan, Longxu Dou, Yunlong Feng, Dingzirui Wang, Yang Xu, Enbo Wang, Qiguang Chen

Abstract

The Myers-Briggs Type Indicator (MBTI) is one of the most influential personality theories reflecting individual differences in thinking, feeling, and behaving. MBTI personality detection has garnered considerable research interest and has evolved significantly over the years. However, this task tends to be overly optimistic, as it currently does not align well with the natural distribution of population personality traits. Specifically, the self-reported labels in existing datasets result in data quality issues and the hard labels fail to capture the full range of population personality distributions. In this paper, we identify the task by constructing MBTIBench, the first manually annotated MBTI personality detection dataset with soft labels, under the guidance of psychologists. Our experimental results confirm that soft labels can provide more benefits to other psychological tasks than hard labels. We highlight the polarized predictions and biases in LLMs as key directions for future research.

BibTeX
@inproceedings{li-etal-2025-large,
    title = "Can Large Language Models Understand You Better? An {MBTI} Personality Detection Dataset Aligned with Population Traits",
    author = "Li, Bohan  and
      Guan, Jiannan  and
      Dou, Longxu  and
      Feng, Yunlong  and
      Wang, Dingzirui  and
      Xu, Yang  and
      Wang, Enbo  and
      Chen, Qiguang  and
      Wang, Bichen  and
      Xu, Xiao  and
      Zhang, Yimeng  and
      Qin, Libo  and
      Zhao, Yanyan  and
      Zhu, Qingfu  and
      Che, Wanxiang",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.339/",
    pages = "5071--5081"
}
Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits · COLING 2025