2025
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
ACL 2025long
The application scope of Large Language Models (LLMs) continues to expand, leading to increasing interest in personalized LLMs that align with human values. However, aligning these models with individual values raises significant safety concerns, as certain values may correlate with harmful informat…