← Search

Guanxu Chen

2 accepted papers

2026

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

ICML 2026poster

Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain unclear. Chain-of-thoughts (CoTs) have been commonly utilized to externalize LLMs' thinking, but this strategy fails to accurately reflect LLMs' thinking process.…

Cited by 0SourceScholar
2026

Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models

ICLR 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) for large language models (LLMs) has achieved remarkable progress in enhancing LLMs’ reasoning capabilities on tasks with clear correctness criteria, such as mathematical reasoning tasks. Several training metrics, such as entropy or response leng…

Cited by 0SourcecodeScholar