← Search

Brian Christian

1 accepted papers

2026

Reward Models Inherit Value Biases from Pretraining

ICLR 2026poster

Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pre-trained and post-trained LLMs themselves. Because RMs are initialized from LLMs, they inherit representations that shape their behavior, but the nature and extent of t…

Cited by 0SourceScholar