2026
MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
AAAI 2026technical
Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers pro