← Search

Qi Cheng

3 accepted papers

2026

Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning

ICLR 2026poster

Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs). Yet on-policy algorithms such as Group Relative Policy Optimization (GRPO) often suffer in early training: noisy gradients from low-quality rollouts lead to unstable updates and inefficient explora…

Cited by 0SourcecodeScholar
2024

Every Answer Matters: Evaluating Commonsense with Probabilistic Measures

ACL 2024long

Large language models have demonstrated impressive performance on commonsense tasks; however, these tasks are often posed as multiple-choice questions, allowing models to exploit systematic biases. Commonsense is also inherently probabilistic with multiple correct answers. The purpose of “boiling wa…