2026
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
ICML 2026poster
Recent advances in reinforcement learning for code generation have made robust environments essential to prevent reward hacking. As LLMs increasingly serve as evaluators in code-based RL, their ability to detect reward hacking remains understudied. In this paper, we propose a novel taxonomy of rewar…