AAAI 2026technical0 citations

Moral Change or Noise? On Problems of Aligning AI with Temporally Unstable Human Feedback

Vijay Keswani, Cyrus Cousins, Breanna Nguyen, Vincent Conitzer, Hoda Heidari, Jana Schaich Borg, Walter Sinnott-Armstrong

Abstract

Alignment methods in moral domains seek to elicit moral preferences of human stakeholders and incorporate them into AI. This presupposes moral preferences as static targets, but such preferences often evolve over time. Proper alignment of AI to dynamic human preferences should ideally account for "legitimate" changes to moral reasoning, while ignoring changes related to attention deficits, cognitive biases, or other arbitrary factors. However, common AI alignment approaches largely neglect temporal changes in preferences, posing serious challenges to proper alignment, especially in high-stakes applications of AI, e.g., in healthcare domains, where misalignment can jeopardize the trustworthiness of the system and yield serious individual and societal harms. This work investigates the extent to which people

BibTeX
@inproceedings{aaai2026_moralchangeornoi,
  title = {Moral Change or Noise? On Problems of Aligning AI with Temporally Unstable Human Feedback},
  author = {Vijay Keswani and Cyrus Cousins and Breanna Nguyen and Vincent Conitzer and Hoda Heidari and Jana Schaich Borg and Walter Sinnott-Armstrong},
  booktitle = {AAAI 2026},
  year = {2026}
}
Moral Change or Noise? On Problems of Aligning AI with Temporally Unstable Human Feedback · AAAI 2026