2025
An Evolutionary Perspective on AI Alignment (Student Abstract)
AAAI 2025technical
Attempting to align AI capabilities and value structures by means of value elicitation from humans, such as through Reinforcement Learning from Human Feedback (RLHF), is a computational challenge that raises both psychological and philosophical questions. Adopting an evolutionary perspective on the…