← Search

Diyuan Wu

4 accepted papers

2026

Improved Scaling Laws via Weak-to-Strong Generalization in Random Features Ridge Regression

ICML 2026poster

It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenomenon of weak-to-strong generalization exemplifies the advantage of this two-stage procedure: a strong student is trained on imperfect labels obtained fr…

Cited by 0SourceScholar
2025

Attention with Trained Embeddings Provably Selects Important Tokens

NeurIPS 2025poster

Token embeddings play a crucial role in language modeling but, despite this practical relevance, their theoretical understanding is limited. Our paper addresses the gap by characterizing the structure of embeddings obtained via gradient descent. Specifically, we consider a one-layer softmax attentio…

Cited by 0SourceScholar
2025

Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime

ICML 2025spotlight

Neural Collapse is a phenomenon where the last-layer representations of a well-trained neural network converge to a highly structured geometry. In this paper, we focus on its first (and most basic) property, known as NC1: the within-class variability vanishes. While prior theoretical studies establ…

Cited by 0SourcePDFScholar
2024

The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information

NeurIPS 2024poster

The rising footprint of machine learning has led to a focus on imposing model sparsity as a means of reducing computational and memory costs. For deep neural networks (DNNs), the state-of-the-art accuracy-vs-sparsity is achieved by heuristics inspired by the classical Optimal Brain Surgeon (OBS) fra…

Cited by 0SourcePDFScholar