2026
ICPO: Provable and Practical In-Context Policy Optimization for Test-Time Scaling
ICLR 2026poster
We study test-time scaling, where a model improves its answer through multi-round self-reflection at inference. We introduce In-Context Policy Optimization (ICPO), in which an agent optimizes its response in context using self-assessed or externally observed rewards without modifying its parameters.…