← Search

Zhihao Cheng

3 accepted papers

2026

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

ICML 2026poster

Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing methods primarily rely on outcome-based supervision to strengthen internal LLM reasoning, often leading to inefficient exploration and sparse rewards. T…

Cited by 0SourceScholar
2023

Offline Quantum Reinforcement Learning in a Conservative Manner

AAAI 2023technical

Recently, to reap the quantum advantage, empowering reinforcement learning (RL) with quantum computing has attracted much attention, which is dubbed as quantum RL (QRL). However, current QRL algorithms employ an online learning scheme, i.e., the policy that is run on a quantum computer needs to inte…

Cited by 8SourcePDFScholar