← Search

Shihong Deng

4 accepted papers

2025

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

EMNLP 2025

Many studies focus on data annotation techniques for training effective PRMs. However, current methods encounter a significant issue when applied to long CoT reasoning processes: they tend to focus solely on the first incorrect step and all preceding steps, assuming that all subsequent steps are inc

Cited by 0SourcePDFScholar
2025

Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning

EMNLP 2025

Recent advances in reasoning language models have witnessed a paradigm shift from short to long CoT pattern. Given the substantial computational cost of rollouts in long CoT models, maximizing the utility of fixed training datasets becomes crucial. Our analysis reveals that negative responses contai

Cited by 0SourcePDFScholar
2020

Potential Driven Reinforcement Learning for Hard Exploration Tasks

IJCAI 2020poster

Experience replay plays a crucial role in Reinforcement Learning (RL), enabling the agent to remember and reuse experience from the past. Most previous methods sample experience transitions using simple heuristics like uniformly sampling or prioritizing those good ones. Since humans can learn from b…

2017

General scale interpolation via context-aware autoregressive model and multiplanar constraint

ICASSP 2017accepted

In this paper, we propose a novel image interpolation algorithm suitable for general scale enlargement. Different from previous AR-based interpolation algorithms which employ predetermined reference configuration to predict pixel values, we consider the context information when building AR models. O…

Cited by 0SourceScholar