2026
Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
ICML 2026poster
Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods focus solely on preference perception accuracies in given specific scenarios, obscuring the critical vulnerabilities of RMs in real-world scenarios. We id…