MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning
Recent advances in multimodal reward modeling have been largely driven by a paradigm shift from discriminative to generative approaches. Building on this progress, recent studies have further employed reinforcement learning with verifiable rewards (RLVR) to enhance multimodal reward models (MRMs). D