← Search

Fuwei Cui

3 accepted papers

2025

An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning

ACL 2025long

Enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) is of great scientific and practical significance. Researchers typically employ process-supervised reward models (PRMs) to guide the reasoning process, effectively improving the models’ reasoning abilities. However, ex…

2025

KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning

NeurIPS 2025poster

Recent advances have demonstrated that integrating reinforcement learning with rule-based rewards can significantly enhance the reasoning capabilities of large language models (LLMs), even without supervised fine-tuning (SFT). However, prevalent reinforcement learning algorithms such as GRPO and its…

Cited by 0SourcecodeScholar
2021

Syntactically Diverse Adversarial Network for Knowledge-Grounded Conversation Generation

EMNLP 2021finding

Generative conversation systems tend to produce meaningless and generic responses, which significantly reduce the user experience. In order to generate informative and diverse responses, recent studies proposed to fuse knowledge to improve informativeness and adopt latent variables to enhance the di…

Cited by 3SourcePDFScholar