2025
Reward Generalization in RLHF: A Topological Perspective
ACL 2025finding
Existing alignment methods share a common topology of information flow, where reward information is collected from humans, modeled with preference learning, and used to tune language models. However, this shared topology has not been systematically characterized, nor have its alternatives been thoro…