← Search

Shiye Lei

2 accepted papers

2026

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

ICML 2026poster

Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing methods primarily rely on outcome-based supervision to strengthen internal LLM reasoning, often leading to inefficient exploration and sparse rewards. T…

Cited by 0SourceScholar