2024
Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch
NeurIPS 2024poster
Detecting and handling misspecified objectives, such as reward functions, has been widely recognized as one of the central challenges within the domain of Artificial Intelligence (AI) safety research. However, even with the recognition of the importance of this problem, we are unaware of any works t…