Contextual Value Alignment
Pierre L. Dognin, Jesus Rios, Ronny Luss, Prasanna Sattigeri, Miao Liu, Inkit Padhi, Matthew Riemer, Manish Nagireddy
Abstract
Developing value-aligned agents is a complex undertaking and an ongoing challenge in the field of AI. Indeed, designing Large Language Models (LLMs) that can balance multiple possibly conflicting moral values based on the context is a problem of paramount importance. In this paper, we propose a system that performs contextual value alignment based on contextual aggregation of possible responses. This aggregation is achieved by integrating a subset of possible LLM responses that are best suited to a user's input while taking into account features extracted about the user's moral preferences. The proposed system trained using the Moral Integrity Corpus displays better alignment to human values than state-of-the-art baselines.
BibTeX
@inproceedings{icassp2025_contextualvaluea,
title = {Contextual Value Alignment},
author = {Pierre L. Dognin and Jesus Rios and Ronny Luss and Prasanna Sattigeri and Miao Liu and Inkit Padhi and Matthew Riemer and Manish Nagireddy and Kush R. Varshney and Djallel Bouneffouf},
booktitle = {ICASSP 2025},
year = {2025}
}