← Search

Guowei Rong

1 accepted papers

2026

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

ICML 2026oral

Reward models learned from human preferences are central to aligning large language models (LLMs) via reinforcement learning from human feedback, yet they are often vulnerable to reward hacking due to noisy annotations and systematic biases such as response length or style. We propose Bayesian Non-N…

Cited by 0SourceScholar