← Search

Jasmine Xinze Li

2 accepted papers

2026

EigenBench: A Comparative Behavioral Measure of Value Alignment

ICLR 2026oral

Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models’ values. Given an ensemble of models, a constitution describing a value system, and…

Cited by 0SourcecodeScholar
2024

ProgressGym: Alignment with a Millennium of Moral Progress

NeurIPS 2024spotlight

Frontier AI systems, including large language models (LLMs), hold increasing influence over the epistemology of human users. Such influence can reinforce prevailing societal values, potentially contributing to the lock-in of misguided moral beliefs and, consequently, the perpetuation of problematic…