2025
Reinforcement Learning with Adaptive Reward Modeling for Expensive-to-Evaluate Systems
ICML 2025poster
Training reinforcement learning (RL) agents requires extensive trials and errors, which becomes prohibitively time-consuming in systems with costly reward evaluations. To address this challenge, we propose adaptive reward modeling (AdaReMo) which accelerates RL training by decomposing the complicate…