2026
Moral Change or Noise? On Problems of Aligning AI with Temporally Unstable Human Feedback
AAAI 2026technical
Alignment methods in moral domains seek to elicit moral preferences of human stakeholders and incorporate them into AI. This presupposes moral preferences as static targets, but such preferences often evolve over time. Proper alignment of AI to dynamic human preferences should ideally account for "l