2026
On The Fragility of Benchmark Contamination Detection in Reasoning Models
ICLR 2026poster
Leaderboards for large reasoning models (LRMs) have turned evaluation into a competition, incentivizing developers to optimize directly on benchmark suites. A shortcut to achieving higher rankings is to incorporate evaluation benchmarks into the training data, thereby yielding inflated performance,…