2026
Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning
AAAI 2026technical
Reinforcement Learning with Verifiable Rewards (RLVR) significantly enhances the reasoning capability of Large Language Models (LLMs). Current RLVR approaches typically conduct training across all generated tokens, but neglect to explore which tokens (e.g., prefix tokens) actually contribute to reas