← Search

Srishti Shekhar Agrawal

1 accepted papers

2024

Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets

NeurIPS 2024spotlight

LLMs are increasingly fine-tuned using RLHF datasets to align them with human preferences and values. However, very limited research has investigated which specific human values are operationalized through these datasets. In this paper, we introduce Value Imprint, a framework for auditing and classi…