2025
Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and Enhancement
ACL 2025long
As Large Language Models (LLMs) advance, aligning them with human values is critical for their responsible development. Value principles serve as the foundation for clarifying alignment goals.Multiple sets of value principles have been proposed, such as HHH (helpful, honest, harmless) and instructio…