← Search

Ruike Song

1 accepted papers

2026

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction

AAAI 2026technical

External reasoning systems combine language models with process reward models (PRMs) to select high-quality reasoning paths for complex tasks such as mathematical problem solving. However, these systems are prone to reward hacking, where high-scoring but logically incorrect paths are assigned high s

Cited by 0SourcePDFScholar