ACL 2025finding0 citations

Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective

Yipeng Kang, Junqi Wang, Yexin Li, Mengmeng Wang, Wenming Tu, Quansen Wang, Hengli Li, Tingjun Wu

Abstract

As large language models (LLMs) become increasingly integrated into critical applications, aligning their behavior with human values presents significant challenges. Current methods, such as Reinforcement Learning from Human Feedback (RLHF), typically focus on a limited set of coarse-grained values and are resource-intensive. Moreover, the correlations between these values remain implicit, leading to unclear explanations for value-steering outcomes. Our work argues that a latent causal value graph underlies the value dimensions of LLMs and that, despite alignment training, this structure remains significantly different from human value systems. We leverage these causal value graphs to guide two lightweight value-steering methods: role-based prompting and sparse autoencoder (SAE) steering, effectively mitigating unexpected side effects. Furthermore, SAE provides a more fine-grained approach to value steering. Experiments on Gemma-2B-IT and Llama3-8B-IT demonstrate the effectiveness and controllability of our methods.

BibTeX
@inproceedings{kang-etal-2025-values,
    title = "Are the Values of {LLM}s Structurally Aligned with Humans? A Causal Perspective",
    author = "Kang, Yipeng  and
      Wang, Junqi  and
      Li, Yexin  and
      Wang, Mengmeng  and
      Tu, Wenming  and
      Wang, Quansen  and
      Li, Hengli  and
      Wu, Tingjun  and
      Feng, Xue  and
      Zhong, Fangwei  and
      Zheng, Zilong",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.1188/",
    doi = "10.18653/v1/2025.findings-acl.1188",
    pages = "23147--23161",
    ISBN = "979-8-89176-256-5"
}
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective · ACL 2025