← Search

Zhenduo Zhang

4 accepted papers

2026

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

ICLR 2026poster

Recent advances in large language models (LLMs) have utilized reinforcement learning with verifiable rewards (RLVR) to improve reasoning capabilities. However, scaling these methods typically requires massive data and extensive rollout computations, leading to high training costs and low data effici…

Cited by 0SourceScholar
2025

Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning

NeurIPS 2025poster

Large reasoning models (LRMs) have demonstrated strong performance on complex reasoning tasks, but often suffer from overthinking, generating redundant content regardless of task difficulty. Inspired by the dual process theory in cognitive science, we propose Adaptive Cognition Policy Optimization (…

Cited by 0SourceScholar
2022

Pose-Invariant Face Recognition via Adaptive Angular Distillation

AAAI 2022technical

Pose-invariant face recognition is a practically useful but challenging task. This paper introduces a novel method to learn pose-invariant feature representation without normalizing profile faces to frontal ones or learning disentangled features. We first design a novel strategy to learn pose-invari…

Cited by 3SourcePDFScholar