2026
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
ICML 2026poster
We propose Re-FORC, an adaptive reward prediction method that, given a context, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, demonstrating improved prediction with longer reasoning a…