2025
Incentivizing LLMs to Self-Verify Their Answers
NeurIPS 2025poster
Large Language Models (LLMs) have demonstrated remarkable progress in complex reasoning tasks through both post-training and test-time scaling laws. While prevalent test-time scaling approaches are often realized by using external reward models to guide the model generation process, we find that onl…