← Search

Lionel Levine

1 accepted papers

2026

EigenBench: A Comparative Behavioral Measure of Value Alignment

ICLR 2026oral

Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models’ values. Given an ensemble of models, a constitution describing a value system, and…

Cited by 0SourcecodeScholar