← Search

Ricardo Olmedo

1 accepted papers

2026

MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model

ICLR 2026poster

With the advent of DeepSeek-R1, a new wave of reinforcement learning (RL) methods has emerged that seem to unlock stronger mathematical reasoning. However, a closer look at the open-source ecosystem reveals a critical limitation: with sufficiently many draws (e.g., $\texttt{pass@1024}$), existing ba…

Cited by 0SourceScholar