← Search

Ziliang Wang

2 accepted papers

2026

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs

ICLR 2026poster

While search-augmented large language models (LLMs) exhibit impressive capabilities, their reliability in complex multi-hop reasoning remains limited. This limitation arises from three fundamental challenges: decomposition errors, where tasks are incorrectly broken down; retrieval missing, where key…

Cited by 0SourceScholar
2025

StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

EMNLP 2025

Efficient multi-hop reasoning requires Large Language Models (LLMs) based agents to acquire high-value external knowledge iteratively. Previous work has explored reinforcement learning (RL) to train LLMs to perform search-based document retrieval, achieving notable improvements in QA performance, bu