← Search

Leila Khalatbari

1 accepted papers

2025

High-Dimension Human Value Representation in Large Language Models

NAACL 2025long

The widespread application of Large Language Models (LLMs) across various tasks and fields has necessitated the alignment of these models with human values and preferences. Given various approaches of human value alignment, such as Reinforcement Learning with Human Feedback (RLHF), constitutional le…