← Search

Yanwei Ren

2 accepted papers

2026

ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling

ICLR 2026poster

Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time scaling (TTS). However, while most PRMs exhibit substantial gains in mathematical domains, the scarcity of domain-specific…

Cited by 0SourceScholar
2025

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation

NeurIPS 2025poster

Enhancing large language models by simply scaling up datasets has begun to yield diminishing returns, shifting the spotlight to data quality. Monte Carlo Tree Search (MCTS) has emerged as a powerful technique for generating high-quality chain-of-thought data, yet conventional approaches typically re…

Cited by 0SourceScholar