← Search

Ruiyu Xiao

3 accepted papers

2026

MRPO: Magnitude-Regularized Policy Optimization via L1 Constraints

ICML 2026poster

Reinforcement learning (RL) for large language models (LLMs) relies on imperfect reward supervision, necessitating constraints on policy updates to prevent overfitting. Nevertheless, the widely adopted KL constraint over-penalizes actions with low reference probabilities and lacks the sparsity to di…

Cited by 0SourceScholar
2025

Stimulate the Critical Thinking of LLMs via Debiasing Discussion

EMNLP 2025

Large language models (LLMs) often succumb to users’ viewpoints when faced with conflicting perspectives. We identify two key biases underlying this issue : stance homogeneity bias and human preference bias. To address these biases, we propose a novel two-stage training framework: Multi-stance Discu

Cited by 0SourcePDFScholar
2024

Prove Your Point!: Bringing Proof-Enhancement Principles to Argumentative Essay Generation

EMNLP 2024main

Argumentative essay generation (AEG) aims to generate complete texts on specific controversial topics or debates. Although current AEG methods can generate individual opinions, they often overlook the high-level connections between these opinions. This often leads to the generated results being mire…

Cited by 0SourcePDFScholar