← Search

Tianpeng Bu

4 accepted papers

2026

Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization

ICML 2026poster

Recent advances in online reinforcement learning (RL) for large language models (LLMs) have demonstrated promising performance in complex reasoning tasks. However, they often exhibit an imbalanced exploration–exploitation trade-off, resulting in unstable optimization and sub-optimal performance. We …

Cited by 0SourceScholar
2026

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

ICML 2026spotlight

While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge this gap at both the evaluation and data levels, we introduce GUI-RobustEval and propose Robustness-driven Trajectory Synthesis. GUI-RobustEval containi…

Cited by 0SourceScholar
2025

Enhanced Data Synthesis for LLM through Reasoning Structures Generated by Hierarchical GFlowNet

ACL 2025finding

Large language models (LLMs) excel in problem-solving but require training data with diverse reasoning processes. Existing methods mainly optimize instruction-response pairs but lack a systematic design for the underlying reasoning structure. This paper proposes RSS: a Reasoning Structure driven dat…

2023

A Dynamic Dual-Processing Object Detection Framework Inspired by the Brain's Recognition Mechanism

ICCV 2023poster

There are two main approaches to object detection: CNN-based and Transformer-based. The former views object detection as a dense local matching problem, while the latter sees it as a sparse global retrieval problem. Research in neuroscience has shown that the recognition decision in the brain is bas…

Cited by 2PDFScholar