← Search

Chenghua He

2 accepted papers

2025

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

EMNLP 2025

Many studies focus on data annotation techniques for training effective PRMs. However, current methods encounter a significant issue when applied to long CoT reasoning processes: they tend to focus solely on the first incorrect step and all preceding steps, assuming that all subsequent steps are inc

Cited by 0SourcePDFScholar
2024

M2RL: A Multi-player Multi-agent Reinforcement Learning Framework for Complex Games

IJCAI 2024poster

Distributed deep reinforcement learning (DDRL) has gained increasing attention due to the emerging requirements for addressing complex games like Go and StarCraft. However, how to effectively and stably train bots with asynchronous and heterogeneous agents cooperation and competition for multiple pl…

Cited by 0SourcePDFScholar