← Search

Hailun Lu

3 accepted papers

2026

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation

ICML 2026poster

Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieved strong results in improving the reasoning capabilities of large language models (LLMs). However, GRPO is prone to advantage collapse, a failure mo…

Cited by 0SourceScholar
2026

Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic Exploration

AAAI 2026technical

Large language models (LLMs) have achieved remarkable success in various natural language processing tasks, yet they remain prone to generating factually incorrect outputs—known as "hallucinations". While recent approaches have shown promise for hallucination detection by repeatedly sampling from LL

Cited by 0SourcePDFScholar
2025

ToolFiVe: Enhancing Tool-Augmented LLMs via Tool Filtering and Verification

ICASSP 2025accepted

Tool-augmented Large Language Models (LLMs) provide a robust theoretical foundation for AI agents, with the generation of reasoning plans being a crucial stage. Previous methods for generating reasoning plans primarily rely on In-Context Learning (ICL) or Supervised Fine-Tuning (SFT). However, metho…

Cited by 0SourceScholar