2025
Deep Value Benchmark: Measuring Whether Models Generalize Deep values or Shallow Preferences
NeurIPS 2025spotlight
We introduce the Deep Value Benchmark (DVB), an evaluation framework that directly tests whether large language models (LLMs) learn fundamental human values or merely surface-level preferences. This distinction is critical for AI alignment: Systems that capture deeper values are likely to generalize…