← Search

Minhao Zou

4 accepted papers

2026

OVLR: Efficient, Scalable, and Robust Training via Output-Level Variance-Reduced Likelihood Ratio

ICML 2026poster

Gradient-based optimization is fundamental to deep learning, yet standard backpropagation (BP) is inherently limited by the requirement of differentiability, rendering it brittle when encountering piecewise-constant objectives with vanishing gradients (e.g., hard 0-1 loss) or black-box feedback. Whi…

Cited by 0SourceScholar
2026

RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-Training

ICLR 2026poster

Reinforcement learning with verifiable reward has recently emerged as a central paradigm for post-training large language models (LLMs); however, prevailing mean-based methods, such as Group Relative Policy Optimization (GRPO), suffer from entropy collapse and limited reasoning gains. We argue that…

Cited by 0SourcecodeScholar
2025

LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-Context QA

ACL 2025finding

Though current long-context large language models (LLMs) have demonstrated impressive capacities in answering various questions based on extensive text, the lack of citations in their responses makes user verification difficult, leading to concerns about their trustworthiness due to the potential ha…

2025

Multi-hop Self-augmented Graph Contrastive Learning for Node Classification

ICASSP 2025accepted

Current Graph Contrastive Learning (GCL) methods primarily focus on adapting data augmentation techniques from Computer Vision (CV) or Natural Language Processing (NLP) domains. These techniques typically involve modifying input data via node sampling, edge perturbation, or graph structure perturbat…

Cited by 0SourceScholar