2025
S2R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
ACL 2025long
Recent studies have demonstrated the effectiveness of LLM test-time scaling. However, existing approaches to incentivize LLMs’ deep thinking abilities generally require large-scale data or significant training efforts. Meanwhile, it remains unclear how to improve the thinking abilities of less power…