2026
R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
ICML 2026poster
Self-play bootstraps LLM reasoning through an iterative Challenger–Solver loop: the Challenger is trained to generate questions that target the Solver's capabilities, and the Solver is optimized on the generated data to expand its reasoning skills. However, existing frameworks like R-Zero often exhi…