2026
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
ICML 2026poster
Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomous systems. We introduce the Reward Hacking Benchmark (RHB), a suite of multi-step tasks requiring sequential tool operations with naturalistic short…