2026
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
ICLR 2026poster
Test-time scaling offers a promising path to improve LLM reasoning by utilizing more compute at inference time; however, the true promise of this paradigm lies in extrapolation (i.e., improvement in performance on hard problems as LLMs keep "thinking" for longer, beyond the maximum token budget they…