← Search

Chunmei Xie

1 accepted papers

2026

UCPO: Uncertainty-Aware Policy Optimization

ICML 2026poster

The key to building trustworthy Large Language Models (LLMs) lies in endowing them with inherent uncertainty expression capabilities to mitigate the hallucinations that restrict their high-stakes applications. However, existing RL paradigms such as GRPO often suffer from Advantage Bias due to binary…

Cited by 0SourceScholar