← Search

Gavin Zhang

3 accepted papers

2026

Reinforcement Learning from Dynamic Critic Feedback for Free-Form Generations

ICLR 2026poster

Open-ended generation tasks require outputs to satisfy diverse and often implicit task-specific evaluation rubrics. The sheer number of relevant rubrics leads to prohibitively high verification costs and incomplete assessments of a response, making reinforcement learning (RL) post-training with rubr…

Cited by 0SourceScholar
2022

Accelerating SGD for Highly Ill-Conditioned Huge-Scale Online Matrix Completion

NeurIPS 2022accept

The matrix completion problem seeks to recover a $d\times d$ ground truth matrix of low rank $r\ll d$ from observations of its individual elements. Real-world matrix completion is often a huge-scale optimization problem, with $d$ so large that even the simplest full-dimension vector operations with…

2021

Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization

NeurIPS 2021poster

In practical instances of nonconvex matrix factorization, the rank of the true solution $r^{\star}$ is often unknown, so the rank $r$ of the model can be over-specified as $r>r^{\star}$. This over-parameterized regime of matrix factorization significantly slows down the convergence of local search a…

Cited by 43SourcePDFScholar