← Search

Zeyu Gan

4 accepted papers

2026

Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents

ICLR 2026poster

Large language model (LLM)–based agents are increasingly trained with reinforcement learning (RL) to enhance their ability to interact with external environments through tool use, particularly in search-based settings that require multi-turn reasoning and knowledge acquisition. However, existing app…

Cited by 0SourcecodeScholar
2025

Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning

ICML 2025poster

Test-time scaling, which is also often referred to as *slow-thinking*, has been demonstrated to enhance multi-step reasoning in large language models (LLMs). However, despite its widespread utilization, the mechanisms underlying slow-thinking methods remain poorly understood. This paper explores the…

2025

Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

ICLR 2025poster

Synthetic data has become a pivotal resource in post-training tasks for large language models (LLMs) due to the scarcity of high-quality, specific data. While various methods have been developed to generate synthetic data, there remains a discernible gap between the practical effects of synthetic da…

2023

Superclass Learning With Representation Enhancement

CVPR 2023poster

In many real scenarios, data are often divided into a handful of artificial super categories in terms of expert knowledge rather than the representations of images. Concretely, a superclass may contain massive and various raw categories, such as refuse sorting. Due to the lack of common semantic fea…

Cited by 5SourcePDFScholar