← Search

Zehong Cao

2 accepted papers

2026

SSVPO: Effective Step-Level Credit Assignment for RL Training of Language Models

ICLR 2026poster

Language models have shown strong performance on mathematical reasoning tasks. Post-training with outcome-based reinforcement learning (RL) can further enhance reasoning but is inefficient because it relies solely on final rewards. Recent credit assignment–based RL methods provide intermediate feedb…

Cited by 0SourceScholar
2026

TrustworthyQENN: A Quantum Evidential Neural Network Based on Complex-Valued Contrastive Learning for Uncertainty Pattern Classification

ICML 2026poster

Out-of-Distribution (OOD) detection requires accurately classifying In-Distribution (ID) samples while effectively distinguishing anomalous OOD data. However, existing methodologies predominantly rely on real-valued magnitude features, neglecting the semantic richness embedded in phase information, …

Cited by 0SourceScholar