CUARewardBench: Benchmark for Evaluating Reward Models on Computer-using Agent Trajectories
Computer-using agents (CUAs) enable task completion through natural interaction with operating systems and software interfaces. While script-based verifiers are widely adopted for evaluation, they suffer from limited scalability and inability to provide step-wise assessment. Reward models offer prom…