ICLR 2026poster0 citations

ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling

Haotian Zhang, Liu Liu, Baosheng Yu, Jiayan Qiu, Likang Xiao, Yanwei Ren, Quan Chen, Xianglong Liu

Abstract

Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time scaling (TTS). However, while most PRMs exhibit substantial gains in mathematical domains, the scarcity of domain-specific training data and knowledge-based learning patterns limits their generalization ability when faced with other domains. To address this limitation, we shift the learning objective from verifying domain-specific knowledge to modeling domain-agnostic logical flow. Centering on \textit{contextual coherence} between chain-of-thought (CoT) steps, our approach is realized through a novel data annotation and training framework, which enhances the model's generalization capabilities across diverse domains. For instance, our resulting model, \textbf{ContextPRM}, achieves a notable 6.5\% average accuracy improvement over the majority voting baseline via weighted majority voting across nine non-mathematical domains in MMLU-Pro, including law, history, and philosophy, significantly surpassing the 2.2\% improvement from VersaPRM and 0.5\% gains from other mathematics-focused PRMs, demonstrating consistent performance across both mathematical and non-mathematical domains.

Process Reward Models
BibTeX
@inproceedings{
zhang2026contextprm,
title={Context{PRM}: Leveraging Contextual Coherence for multi-domain Test-Time Scaling},
author={Haotian Zhang and Liu Liu and Baosheng Yu and Jiayan Qiu and Likang Xiao and Yanwei Ren and Quan Chen and Xianglong Liu},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=9H0gBsNjCv}
}
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling · ICLR 2026