2026
SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
ICML 2026poster
Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-training Large Language Models (LLMs) remains underexplored. We propose Self-Reflective Policy Optimization (SRPO), a framewor…