← Search

Ruiyang Zhou

2 accepted papers

2025

ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning

NeurIPS 2025poster

Recent advances in large language models have been driven by reinforcement learning (RL)-style post-training, which improves reasoning by optimizing model outputs based on reward or preference signals. GRPO-style approaches implement this by using self-generated samples labeled by an outcome-based v…

Cited by 0SourceScholar
2024

Is LLM a Reliable Reviewer? A Comprehensive Evaluation of LLM on Automatic Paper Reviewing Tasks

COLING 2024main

The use of large language models (LLM), especially ChatGPT, to help with research has come into practice. Researchers use it for timely advice and hope to obtain in-depth feedback. However, can LLM be a qualified and reliable reviewer? Although there already exist several review-related datasets, fe…

Cited by 40SourcePDFScholar