2026
Factored Causal Representation Learning for Robust Reward Modeling in RLHF
ICML 2026poster
A reliable reward model is essential for aligning large language models (LLMs) with human preferences through reinforcement learning from human feedback (RLHF). However, standard reward models are susceptible to spurious features that are not causally related to human labels. This can lead to *rewar…