← Search

Qining Zhang

3 accepted papers

2025

Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference

ICLR 2025poster

Reward inference (learning a reward model from human preferences) is a critical intermediate step in the Reinforcement Learning from Human Feedback (RLHF) pipeline for fine-tuning Large Language Models (LLMs). In practice, RLHF faces fundamental challenges such as distribution shift, reward model ov…

Cited by 1SourcePDFScholar
2024

Deep Reinforcement Learning for Early Diagnosis of Lung Cancer

AAAI 2024technical

Lung cancer remains the leading cause of cancer-related death worldwide, and early diagnosis of lung cancer is critical for improving the survival rate of patients. Performing annual low-dose computed tomography (LDCT) screening among high-risk populations is the primary approach for early diagnosis…

2023

Fast and Regret Optimal Best Arm Identification: Fundamental Limits and Low-Complexity Algorithms

NeurIPS 2023poster

This paper considers a stochastic Multi-Armed Bandit (MAB) problem with dual objectives: (i) quick identification and commitment to the optimal arm, and (ii) reward maximization throughout a sequence of $T$ consecutive rounds. Though each objective has been individually well-studied, i.e., best arm…

Cited by 9SourcePDFScholar