← Search

Shaoning Sun

2 accepted papers

2026

Reward Modeling from Natural Language Human Feedback

ICML 2026poster

Reinforcement Learning with Verifiable Reward (RLVR) on preference data has become the mainstream approach for training Generative Reward Models (GRMs). Typically, GRMs generate reasoning chains ending with critiques and preference labels, with RLVR using label correctness as the training reward. Ho…

Cited by 0SourceScholar
2025

Improve LLM-as-a-Judge Ability as a General Ability

EMNLP 2025

LLM-as-a-Judge leverages the generative and reasoning capabilities of large language models (LLMs) to evaluate LLM responses across diverse scenarios, providing accurate preference signals. This approach plays a vital role in aligning LLMs with human values. Recent studies have raised many methods t

Cited by 0SourcePDFScholar