← Search

Siyao Song

2 accepted papers

2026

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

ICML 2026poster

Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing methods primarily rely on outcome-based supervision to strengthen internal LLM reasoning, often leading to inefficient exploration and sparse rewards. T…

Cited by 0SourceScholar
2026

FIRE: Learning to Navigate and Act on Real-World Files via Stateful Reinforcement Learning

ICML 2026poster

Large language models still struggle to reliably answer questions grounded in real-world files like spreadsheets and slides, where evidence is scattered across irregular layouts and heterogeneous formats. We address this by formalizing File Reasoning, a setting where agents must interact directly wi…

Cited by 0SourceScholar